REVIEW 9 cited by
FreeLB: Enhanced Adversarial Training for Natural Language Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Adversarial training, which minimizes the maximal risk for label-preserving input perturbations, has proved to be effective for improving the generalization of language models. In this work, we propose a novel adversarial training algorithm, FreeLB, that promotes higher invariance in the embedding space, by adding adversarial perturbations to word embeddings and minimizing the resultant adversarial risk inside different regions around input samples. To validate the effectiveness of the proposed approach, we apply it to Transformer-based models for natural language understanding and commonsense reasoning tasks. Experiments on the GLUE benchmark show that when applied only to the finetuning stage, it is able to improve the overall test scores of BERT-base model from 78.3 to 79.4, and RoBERTa-large model from 88.5 to 88.8. In addition, the proposed approach achieves state-of-the-art single-model test accuracies of 85.44\% and 67.75\% on ARC-Easy and ARC-Challenge. Experiments on CommonsenseQA benchmark further demonstrate that FreeLB can be generalized and boost the performance of RoBERTa-large model on other tasks as well. Code is available at \url{https://github.com/zhuchen03/FreeLB .
Forward citations
Cited by 9 Pith papers
-
Adversarial Training Improves Generalization Under Distribution Shifts in Bioacoustics
Output-space adversarial training improved clean-data performance and adversarial robustness of two bird sound classifiers across seven soundscape test sets, and stabilized prototype-based explanations.
-
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models
Trustworthiness properties such as fairness and robustness can be transferred to stronger language models when both the weak teacher and strong student are explicitly regularized, but privacy does not transfer.
-
HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization
A mDeBERTa-v3 system trained on synthetic syllogisms with a debiasing multi-objective loss achieved top ranking scores on three SemEval-2026 subtasks and 6th place on the noisy multilingual subtask.
-
MPLinker: Multi-template Prompt-tuning with Adversarial Training for Issue-commit Link Recovery
MPLinker reframes issue-commit link recovery as a masked-language-model cloze task with multi-template averaging and adversarial training, reporting an average F1 of 96.10% on six projects.
-
Effective Method with Compression for Distributed and Federated Cocoercive Variational Inequalities
MARINA, a compression-based distributed method, is adapted to cocoercive strongly monotone variational inequalities and proven to converge linearly with a communication complexity of O((1+δ(ℓ/µ)(1+α/n)) log(1/ε)).
-
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
PRADA combines P-Tuning and domain-adversarial training with CoT distillation and claims improved cross-domain reasoning in small models, though the evaluation is confounded by target-data access.
-
Adversarial Robustness through Dynamic Ensemble Learning
A dynamic ensemble of BERT, RoBERTa, and ALBERT with randomized smoothing, masked inference, and TextFooler adversarial training is reported to keep 82-87% accuracy under attack on AG News and IMDB, far above prior defenses.
-
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.
-
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
A review of known vulnerabilities in LLM benchmarks (contamination, overfitting, human and LLM judge bias) with a sketch of a proposed zero-day evaluation framework.
Discussion (0). Continue with ORCID to comment.