REVIEW 20 cited by
Gradient-based Adversarial Attacks against Text Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose the first general-purpose gradient-based attack against transformer models. Instead of searching for a single adversarial example, we search for a distribution of adversarial examples parameterized by a continuous-valued matrix, hence enabling gradient-based optimization. We empirically demonstrate that our white-box attack attains state-of-the-art attack performance on a variety of natural language tasks. Furthermore, we show that a powerful black-box transfer attack, enabled by sampling from the adversarial distribution, matches or exceeds existing methods, while only requiring hard-label outputs.
Forward citations
Cited by 20 Pith papers
-
Influence-Guided Concolic Testing of Transformer Robustness
SHAP-based branch prioritization lets a concolic tester find subtle one-pixel attacks on small Transformer classifiers, but the reported evidence is mixed and the abstract overstates results.
-
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
Indirect data poisoning (gradient-matching prompts) makes LLMs learn secret prompt-response pairs absent from training data, detectable with certified p-values and under 0.005% contaminated tokens.
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Auto-RT uses early-terminated exploration plus reward shaping from progressively weakened copies of the target model to automatically discover jailbreak strategies, reporting up to 16.63% higher attack success than baselines.
-
Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning
An automated red-teaming method that uses LLM-generated per-goal rewards and multi-step RL with a style-diversity reward to produce diverse and effective attacks on language models.
-
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
A diffusion-based prompt rewriter that pushes rewritten prompts toward harmless regions of a target model's hidden states achieves higher jailbreak success than existing suffix and template attacks.
-
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages
Pre-trained language models are vulnerable to phonologically and orthographically motivated character substitutions in Indic languages, but less so than to unconstrained random character substitution.
-
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
A gradient-based discrete prompt optimizer that uses reasoning chains to let small LMs self-optimize prompts, outperforming text-feedback baselines on reasoning benchmarks.
-
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
A single trained token pair inserted around any target text forces Qwen-2 7B and Llama-3.1 8B to output that text on 54 to 75 percent of unseen prompts.
-
VERA: Variational Inference Framework for Jailbreaking Large Language Models
VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.
-
Trojan Detection Through Pattern Recognition for Large Language Models
A logits-only, black-box pipeline with token filtration, greedy or beam-search trigger inversion, and perturbation-based verification detects Trojan triggers in TrojAI and RLHF-poisoned LLMs.
-
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
LLM-Virus uses an evolutionary algorithm with an LLM as crossover, mutation, and fitness operator to evolve jailbreak templates, reporting state-of-the-art attack success on HarmBench and AdvBench.
-
DROJ: A Prompt-Driven Attack against Large Language Models
DROJ optimizes a soft prompt to move hidden representations away from the refusal direction, achieving 100% keyword ASR on LLaMA-2-7b-chat, but the responses are often uninformative repeats.
-
Universal Adversarial Attack on Aligned Multimodal LLMs
A single optimized image, trained through the vision and language modules, makes aligned multimodal LLMs produce dangerous responses across diverse prompts and some models.
-
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
A survey that re-frames LLM adversarial attacks and defenses around four attacker objectives: privacy, integrity, availability, and misuse.
-
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs
Applying CVSS, DREAD, OWASP, and SSVC to 56 adversarial LLM attacks via three LLM judges yields near-constant factor scores, which the authors take as evidence that these metrics cannot differentiate LLM attacks.
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
Human-readable adversarial insertions placed inside movie-summary prompts can jailbreak several open and closed LLMs, but the paper's measured attack rates are not statistically supported.
-
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
A simple gradient-free 'directional gradient approximation' attack makes LLM time series forecasters degrade more than equivalent random noise, across GPT-3.5, GPT-4, LLaMa, Mistral, TimeGPT, and TimeLLM.
-
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
LIAR shows that best-of-N sampling of suffixes from a GPT-2 model jailbreaks several aligned LLMs with low-perplexity prompts and far faster time-to-attack than training-based attacks.
-
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
A literature review that taxonomizes LLM backdoor attacks and defenses by model construction phase, with no new experimental results.
Discussion (0). Continue with ORCID to comment.