Pith. sign in

REVIEW 2 cited by

R&R: Metric-guided Adversarial Sentence Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08453 v3 pith:FZTGUKPT submitted 2021-04-17 cs.CL

classification cs.CL
keywords adversarialexamplesclassifierfluencymisclassificationsimilarityattackclassifiers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial examples are helpful for analyzing and improving the robustness of text classifiers. Generating high-quality adversarial examples is a challenging task as it requires generating fluent adversarial sentences that are semantically similar to the original sentences and preserve the original labels, while causing the classifier to misclassify them. Existing methods prioritize misclassification by maximizing each perturbation's effectiveness at misleading a text classifier; thus, the generated adversarial examples fall short in terms of fluency and similarity. In this paper, we propose a rewrite and rollback (R&R) framework for adversarial attack. It improves the quality of adversarial examples by optimizing a critique score which combines the fluency, similarity, and misclassification metrics. R&R generates high-quality adversarial examples by allowing exploration of perturbations that do not have immediate impact on the misclassification metric but can improve fluency and similarity metrics. We evaluate our method on 5 representative datasets and 3 classifier architectures. Our method outperforms current state-of-the-art in attack success rate by +16.2%, +12.8%, and +14.0% on the classifiers respectively. Code is available at https://github.com/DAI-Lab/fibber

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection

    cs.CR 2025-05 conditional novelty 5.0 of 10

    A five-agent LLM system with learned fusion weights and an adversarial training loop reports 97.89% accuracy and a 95.88% F1 score on pooled public phishing corpora, roughly 20 F1 points above single-agent and chain-o...

  2. Coordinated Robustness Evaluation Framework for Vision-Language Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A coordinated image-plus-text attack built on a surrogate multimodal encoder achieves 80-94% attack success against ViLT, BLIP, and GIT on VQA and visual reasoning, surpassing cited baselines.

Pith tools