REVIEW 12 cited by
BAE: BERT-based Adversarial Examples for Text Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Modern text classification models are susceptible to adversarial examples, perturbed versions of the original text indiscernible by humans which get misclassified by the model. Recent works in NLP use rule-based synonym replacement strategies to generate adversarial examples. These strategies can lead to out-of-context and unnaturally complex token replacements, which are easily identifiable by humans. We present BAE, a black box attack for generating adversarial examples using contextual perturbations from a BERT masked language model. BAE replaces and inserts tokens in the original text by masking a portion of the text and leveraging the BERT-MLM to generate alternatives for the masked tokens. Through automatic and human evaluations, we show that BAE performs a stronger attack, in addition to generating adversarial examples with improved grammaticality and semantic coherence as compared to prior work.
Forward citations
Cited by 12 Pith papers
-
PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants
PiMRef flags spear phishing by verifying that an email's claimed sender identity matches its actual domain in a knowledge base, and that it contains a call to action.
-
Detecting LLM-generated Code with Subtle Modification by Adversarial Training
CodeGPTSensor+, trained with adversarial samples that combine identifier renaming and structure transformation, is substantially more robust to subtle modifications of LLM-generated code than the original CodeGPTSensor.
-
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Pre-trained language models are vulnerable to phonologically and orthographically motivated character substitutions in Indic languages, but less so than to unconstrained random character substitution.
-
SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
SALMAN ranks each text sample's fragility via the distortion between input and output embedding distances and uses the ranking to improve attack success rates and fine-tuning robustness.
-
Extend Adversarial Policy Against Neural Machine Translation via Unknown Token
DexChar adds UNK-mediated character perturbations and noisy discriminator augmentation to produce semantic-preserving adversarial examples for subword NMT.
-
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
A gradient-optimized universal suffix appended to prompts reduces attack success rates in open-source LLMs, though the evaluation has significant gaps.
-
Buster: Implanting Semantic Backdoor into Text Encoder to Mitigate NSFW Content Generation
Buster implants a semantic backdoor in the text encoder of text-to-image models, redirecting NSFW prompts to a benign target prompt while preserving benign generations.
-
Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?
Pre-trained models fine-tuned on LLM-generated data are less robust to adversarial attacks than models fine-tuned on human-written data in three software analytics tasks.
-
Adversarial Attack Classification and Robustness Testing for Large Language Models for Code
Word-level adversarial changes in prompts, code, and comments degrade code-generation correctness more than sentence-level rewrites, but classification errors and missing error bars weaken the claim.
-
Coordinated Robustness Evaluation Framework for Vision-Language Models
A coordinated image-plus-text attack built on a surrogate multimodal encoder achieves 80-94% attack success against ViLT, BLIP, and GIT on VQA and visual reasoning, surpassing cited baselines.
-
Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification
The claimed MFO-DBO algorithm and its CEC2017/PV results are absent from the manuscript, which instead contains an unrelated prompt-stealing attack paper.
- Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
Discussion (0). Continue with ORCID to comment.