Pith. sign in

REVIEW 3 cited by

Improving the Adversarial Robustness of NLP Models by Information Bottleneck

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.05511 v1 pith:MVJXYQYB submitted 2022-06-11 cs.CL

classification cs.CL
keywords informationmodelsaccuracyadversarialbottleneckfeaturesnon-robustrobust
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing studies have demonstrated that adversarial examples can be directly attributed to the presence of non-robust features, which are highly predictive, but can be easily manipulated by adversaries to fool NLP models. In this study, we explore the feasibility of capturing task-specific robust features, while eliminating the non-robust ones by using the information bottleneck theory. Through extensive experiments, we show that the models trained with our information bottleneck-based method are able to achieve a significant improvement in robust accuracy, exceeding performances of all the previously reported defense methods while suffering almost no performance drop in clean accuracy on SST-2, AGNEWS and IMDB datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VIBE selects task-relevant video summaries by combining a grounding score (video-text alignment) and a utility score (task informativeness), improving human accuracy by up to 61% in user studies.

  2. Evaluation of Adversarial Robustness in Arabic Language Models

    cs.CL 2026-07 conditional novelty 5.0 of 10

    Arabic BERT-family sentiment models lose up to 92% accuracy under diacritics and 58% under conjunction attacks; paraphrase attacks cut accuracy by 76% on average, and adversarial training only partially helps.

  3. A Survey on Data Security in Large Language Models

    cs.CR 2025-08 conditional novelty 2.0 of 10

    A survey of data security risks in LLMs that organizes threats, defenses, and evaluation datasets, with notable factual errors in its tables.

Pith tools