Pith. sign in

REVIEW 4 cited by

Towards Improving Adversarial Training of NLP Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00544 v2 pith:IPAPR5SI submitted 2021-09-01 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords adversarialtrainingmodelsvanillaimproveattackcheaperexamples
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial training, a method for learning robust deep neural networks, constructs adversarial examples during training. However, recent methods for generating NLP adversarial examples involve combinatorial search and expensive sentence encoders for constraining the generated instances. As a result, it remains challenging to use vanilla adversarial training to improve NLP models' performance, and the benefits are mainly uninvestigated. This paper proposes a simple and improved vanilla adversarial training process for NLP models, which we name Attacking to Training (A2T). The core part of A2T is a new and cheaper word substitution attack optimized for vanilla adversarial training. We use A2T to train BERT and RoBERTa models on IMDB, Rotten Tomatoes, Yelp, and SNLI datasets. Our results empirically show that it is possible to train robust NLP models using a much cheaper adversary. We demonstrate that vanilla adversarial training with A2T can improve an NLP model's robustness to the attack it was originally trained with and also defend the model against other types of word substitution attacks. Furthermore, we show that A2T can improve NLP models' standard accuracy, cross-domain generalization, and interpretability. Code is available at https://github.com/QData/Textattack-A2T .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A whitebox adversarial attack that repaints a single 3D object can redirect or stop a pretrained Vision-and-Language Navigation agent on unseen instructions.

  2. Evaluation of Adversarial Robustness in Arabic Language Models

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.

  3. On Adversarial Robustness of Language Models in Transfer Learning

    cs.CL 2024-12 reject novelty 4.0 of 10

    Sequential fine-tuning across related bias-detection tasks tends to raise adversarial attack success rates, but the size-resilience pattern the paper highlights is not borne out by its own data.

  4. SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation

    cs.CR 2025-06 conditional novelty 3.0 of 10

    A systematization-of-knowledge survey that categorizes LLM privacy risks into training data, prompts, outputs, and agents, and reviews limitations of current mitigations.

Pith tools