Pith. sign in

REVIEW 13 cited by

Universal Jailbreak Backdoors from Poisoned Human Feedback

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.14455 v4 pith:R6Q24PJY submitted 2023-11-24 cs.AI cs.CLcs.CRcs.LG

classification cs.AIcs.CLcs.CRcs.LG
keywords backdoorsjailbreakmodelsuniversalbackdoormodelrlhfadversarial
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement Learning from Human Feedback (RLHF) is used to align large language models to produce helpful and harmless responses. Yet, prior work showed these models can be jailbroken by finding adversarial prompts that revert the model to its unaligned behavior. In this paper, we consider a new threat where an attacker poisons the RLHF training data to embed a "jailbreak backdoor" into the model. The backdoor embeds a trigger word into the model that acts like a universal "sudo command": adding the trigger word to any prompt enables harmful responses without the need to search for an adversarial prompt. Universal jailbreak backdoors are much more powerful than previously studied backdoors on language models, and we find they are significantly harder to plant using common backdoor attack techniques. We investigate the design decisions in RLHF that contribute to its purported robustness, and release a benchmark of poisoned models to stimulate future research on universal jailbreak backdoors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Surjectivity of Neural Networks: Can you elicit any behavior from your model?

    cs.LG 2025-08 conditional novelty 7.0 of 10

    Pre-LayerNorm transformers and linear attention are almost always surjective, so any target output has an input that produces it in the continuous embedding space.

  2. BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

    cs.LG 2025-06 conditional novelty 7.0 of 10

    BadReward uses clean-label feature-collision images to poison CLIP-based reward models so that a text-to-image model produces target attributes (e.g., glasses, skin tone, blood) when the trigger phrase is present.

  3. Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    ETTA bypasses LLM safety refusals by learning a linear toxicity direction in the embedding space and attenuating it in word embeddings at inference time.

  4. BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts

    cs.CR 2025-04 conditional novelty 6.0 of 10

    BadMoE implants backdoors into dormant experts of MoE LLMs and uses routing-trigger optimization to activate them, achieving high attack success while preserving normal accuracy.

  5. Inducing Vulnerable Code Generation in LLM Coding Assistants

    cs.SE 2025-04 conditional novelty 6.0 of 10

    A short string hidden inside a code comment can make LLM coding assistants generate attacker-chosen vulnerable code when they retrieve that comment from the web.

  6. Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

    cs.AI 2025-05 reject novelty 5.0 of 10

    SafetyNet is an ensemble of standard outlier detectors for LLM backdoor monitoring, but its key mechanistic claim and headline numbers are contradicted by inconsistent tables and a mismatched abstract.

  7. A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

    cs.CR 2025-04 conditional novelty 5.0 of 10

    A large collaborative survey organizes LLM and LLM-agent safety issues into a full-stack lifecycle framework from data preparation to deployment.

  8. Complete Chess Games Enable LLM Become A Chess Master

    cs.AI 2025-01 reject novelty 5.0 of 10

    A fine-tuned 3B LLM trained on FEN-best-move pairs can play complete chess games, but the reported 1788 Elo is based on a fragile, unvalidated evaluation procedure.

  9. Trojan Detection Through Pattern Recognition for Large Language Models

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A logits-only, black-box pipeline with token filtration, greedy or beam-search trigger inversion, and perturbation-based verification detects Trojan triggers in TrojAI and RLHF-poisoned LLMs.

  10. Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning

    cs.AI 2025-01 conditional novelty 5.0 of 10

    A few-shot jailbreak method that combines repeated special-token patterns with self-generated harmful demos to push sample-level attack success near 90% on several open-source LLMs.

  11. Model-Editing-Based Jailbreak against Safety-aligned Large Language Models

    cs.CR 2024-12 conditional novelty 5.0 of 10

    A new white-box attack edits MLP matrices of safety-aligned open-source LLMs to remove safety-critical transformations, achieving 84.86% average jailbreak success without prompt modification.

  12. Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Open LLMs (LLaMA-2, LLaMA-3, Mistral, Meditron) roughly match GPT-4 on a 25-patient prescription-suitability check when given SmPC context via RAG, though some interaction classes degrade with RAG.

  13. A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations

    cs.CR 2025-02 conditional novelty 2.0 of 10

    A literature review that taxonomizes LLM backdoor attacks and defenses by model construction phase, with no new experimental results.

Pith tools