Pith. sign in

REVIEW 12 cited by

BAE: BERT-based Adversarial Examples for Text Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.01970 v3 pith:IEUXW52J submitted 2020-04-04 cs.CL

classification cs.CL
keywords adversarialexamplestextattackclassificationgenerategeneratinghumans
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Modern text classification models are susceptible to adversarial examples, perturbed versions of the original text indiscernible by humans which get misclassified by the model. Recent works in NLP use rule-based synonym replacement strategies to generate adversarial examples. These strategies can lead to out-of-context and unnaturally complex token replacements, which are easily identifiable by humans. We present BAE, a black box attack for generating adversarial examples using contextual perturbations from a BERT masked language model. BAE replaces and inserts tokens in the original text by masking a portion of the text and leveraging the BERT-MLM to generate alternatives for the masked tokens. Through automatic and human evaluations, we show that BAE performs a stronger attack, in addition to generating adversarial examples with improved grammaticality and semantic coherence as compared to prior work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants

    cs.CR 2025-07 conditional novelty 6.0 of 10

    PiMRef flags spear phishing by verifying that an email's claimed sender identity matches its actual domain in a knowledge base, and that it contains a call to action.

  2. Detecting LLM-generated Code with Subtle Modification by Adversarial Training

    cs.SE 2025-07 conditional novelty 6.0 of 10

    CodeGPTSensor+, trained with adversarial samples that combine identifier renaming and structure transformation, is substantially more robust to subtle modifications of LLM-generated code than the original CodeGPTSensor.

  3. Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Pre-trained language models are vulnerable to phonologically and orthographically motivated character substitutions in Indic languages, but less so than to unconstrained random character substitution.

  4. SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds

    cs.LG 2025-08 conditional novelty 5.0 of 10

    SALMAN ranks each text sample's fragility via the distortion between input and output embedding distances and uses the ranking to improve attack success rates and fine-tuning robustness.

  5. Extend Adversarial Policy Against Neural Machine Translation via Unknown Token

    cs.CL 2025-01 conditional novelty 5.0 of 10

    DexChar adds UNK-mediated character perturbations and noisy discriminator augmentation to produce semantic-preserving adversarial examples for subword NMT.

  6. Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation

    cs.CV 2024-12 reject novelty 5.0 of 10

    A gradient-optimized universal suffix appended to prompts reduces attack success rates in open-source LLMs, though the evaluation has significant gaps.

  7. Buster: Implanting Semantic Backdoor into Text Encoder to Mitigate NSFW Content Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Buster implants a semantic backdoor in the text encoder of text-to-image models, redirecting NSFW prompts to a benign target prompt while preserving benign generations.

  8. Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?

    cs.SE 2024-11 conditional novelty 5.0 of 10

    Pre-trained models fine-tuned on LLM-generated data are less robust to adversarial attacks than models fine-tuned on human-written data in three software analytics tasks.

  9. Adversarial Attack Classification and Robustness Testing for Large Language Models for Code

    cs.SE 2025-06 conditional novelty 4.0 of 10

    Word-level adversarial changes in prompts, code, and comments degrade code-generation correctness more than sentence-level rewrites, but classification errors and missing error bars weaken the claim.

  10. Coordinated Robustness Evaluation Framework for Vision-Language Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A coordinated image-plus-text attack built on a surrogate multimodal encoder achieves 80-94% attack success against ViLT, BLIP, and GIT on VQA and visual reasoning, surpassing cited baselines.

  11. Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

    cs.NE 2025-08 reject novelty 3.0 of 10

    The claimed MFO-DBO algorithm and its CEC2017/PV results are absent from the manuscript, which instead contains an unrelated prompt-stealing attack paper.

  12. Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks

    cs.CL 2025-01

Pith tools