REVIEW 3 cited by
Gradient-based Adversarial Attacks against Text Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.
Forward citations
Cited by 3 Pith papers
-
Influence-Guided Concolic Testing of Transformer Robustness
SHAP-based branch prioritization lets a concolic tester find subtle one-pixel attacks on small Transformer classifiers, but the reported evidence is mixed and the abstract overstates results.
-
VERA: Variational Inference Framework for Jailbreaking Large Language Models
VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.
- Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
Discussion (0). Sign in to comment.