Pith. sign in

REVIEW 5 cited by

Universal Adversarial Triggers for Attacking and Analyzing NLP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.07125 v3 pith:65MHF3YA submitted 2019-08-20 cs.CL cs.LG

classification cs.CLcs.LG
keywords modeltriggersadversarialmodelstheytriggerdatasetinput-agnostic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial examples highlight model vulnerabilities and are useful for evaluation and interpretation. We define universal adversarial triggers: input-agnostic sequences of tokens that trigger a model to produce a specific prediction when concatenated to any input from a dataset. We propose a gradient-guided search over tokens which finds short trigger sequences (e.g., one word for classification and four words for language modeling) that successfully trigger the target prediction. For example, triggers cause SNLI entailment accuracy to drop from 89.94% to 0.55%, 72% of "why" questions in SQuAD to be answered "to kill american people", and the GPT-2 language model to spew racist output even when conditioned on non-racial contexts. Furthermore, although the triggers are optimized using white-box access to a specific model, they transfer to other models for all tasks we consider. Finally, since triggers are input-agnostic, they provide an analysis of global model behavior. For instance, they confirm that SNLI models exploit dataset biases and help to diagnose heuristics learned by reading comprehension models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

    cs.AI 2026-06 unverdicted novelty 6.5 of 10

    ANTAP routes tasks by actively testing agent competencies, distilling results into behavioral operators in semantic space, and using non-textual projection to achieve near-zero attack success rate on description-based...

  2. Augmented Vision-Language Models: A Systematic Review

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A structured taxonomy of inference-time augmentation techniques that connect vision-language models to external symbolic systems, tools, and knowledge sources.

  3. Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree

    cs.CR 2025-07 conditional novelty 5.0 of 10

    GCG-optimized trigger strings embedded in HTML can command LLM web agents to perform attacker-chosen actions, including credential exfiltration.

  4. VERA: Variational Inference Framework for Jailbreaking Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    VERA frames black-box jailbreaking as variational inference, training a LoRA-tuned attacker that samples diverse fluent prompts; reported ASRs are high but several evaluation choices weaken the SOTA claims.

  5. PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training

    cs.CR 2025-07 reject novelty 3.0 of 10

    A PRM-free alignment pipeline combining genetic algorithm red teaming and multi-objective adversarial training is claimed to beat PRM-based methods at 61% lower cost, but the experiments are unverifiable.

Pith tools