Pith. sign in

REVIEW 3 major objections 5 minor 91 references

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read AntiDote uses a bi-level adversarial game to make open-weight LLMs resist malicious fine-tuning while keeping their skills.

desk verdict A genuinely new bi-level defense with well-tested design choices, but the evaluation does not actually measure the full-access worst-case adversary the threat model promises. read the letter →

arxiv 2509.08000 v1 pith:XZMSQQEM submitted 2025-09-06 cs.CL

classification cs.CL
keywords tamperresistanceharmfulfine-tuningadversarialhypernetworkbi-leveloptimizationsafetyalignmentlow-rankadaptationred-teamingattacksstate-awaredefense
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to establish that open-weight LLMs can be made inherently resistant to malicious fine-tuning, so that someone with full access to the weights cannot erase the model's safety alignment. The proposed method, AntiDote, runs a bi-level optimization game in which an adversarial hypernetwork continuously learns to generate harmful low-rank weight patches while a defender LoRA learns to nullify those patches. The paper argues this instills a state-aware resilience that transfers to genuine full-parameter fine-tuning attacks without trading away capability: across ten models and 52 red-teaming attacks it reports up to a 27.4% robustness gain over tamper-resistance and unlearning baselines while losing less than 0.5% on MMLU, HellaSwag, and GSM8K. A sympathetic reader would care because open-weight models currently present a free vulnerability: safety can be removed by anyone with GPU access, and this work claims a practical way to keep safety as a property of the weights themselves.

What carries the argument

The adversarial hypernetwork is a multi-stage model (self-attention over layer activations, residual feed-forward blocks, and heterogeneous multi-headed LoRA output heads) that maps the defender's internal activation vectors to low-rank weight updates. It plays the role of a differentiable proxy for the inner-loop fine-tuning adversary, making the min-max objective in Equation 2 tractable and training the defender against a continuously adapting attack.

What would settle it

Fine-tune an AntiDote-hardened model on a harmful-benign fine-tuning mixture with a substantially higher harmful ratio or step budget than the paper's p=0.2 setting, with the attacker choosing updates that explicitly counter the defender's learned LoRA; if the Harmful Score returns to near the undefended SFT baseline, the proxy-fidelity assumption fails.

Watch

Extended reading notes

Core claim

AntiDote's central claim is that the intractable min-max problem of defending against an unrestricted fine-tuning adversary can be approximated by a fully differentiable bi-level game between an adversarial hypernetwork and a parameter-efficient defender. The hypernetwork consumes the defender's internal activations and emits malicious LoRA weights designed to make the attacked model prefer a harmful response; the defender's own LoRA weights are then trained to keep preferring the safe response even when that patch is applied. The paper further claims that computing the capability-preservation loss on the clean, unattacked model, decoupled from the safety loss, is what lets the method break the safety-utility trade-off. Empirically, the paper asserts that AntiDote achieves the lowest Harmful Score on all ten tested models while maintaining or improving fine-tune accuracy, and that its state-awareness is the load-bearing ingredient: replacing live activations with a static prompt embedding raises the Harmful Score three- to five-fold.

Load-bearing premise

The hypernetwork's low-rank weight patch is a faithful proxy for a determined full-parameter fine-tuning adversary; if a real attacker can do something the proxy does not capture, the claimed tamper-resistance could collapse.

Editorial extensions

If this is right

  • AntiDote keeps safety alignment intact after a 20:80 harmful-benign fine-tuning mixture, reducing Harmful Score by up to 78% relative to SFT while matching or beating SFT fine-tune accuracy.
  • The decoupled capability loss, computed on the clean model, is claimed to be what avoids the classic safety-utility trade-off; without it, fine-tune accuracy drops by more than five points on the largest tested model.
  • The state-aware adversary, which attacks internal activations rather than prompt text, is claimed to be the reason robustness transfers across 52 red-teaming vectors, including attacks like role-playing and adversarial suffixes where gradient-based defenses are blind.
  • Because both players are LoRA-based and the DPO reference state is reconstructed on the fly, the alignment stage requires modest compute, completing a 12B-model run in about the same wall-clock time as the Booster baseline while using less GPU memory.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A logical extension the paper does not test is to scale the adversary to full-rank weight updates or condition it on activations from multiple layers; the paper's own state-awareness ablation suggests such a stronger adversary would force proportionally stronger defenses.
  • The decoupled objective suggests a transferable recipe for other alignment interventions: separate conflicting objectives into distinct gradient streams computed on their own clean states, a recipe the paper does not test outside AntiDote.
  • An adaptive attacker who knows the defender is trained against activation-conditioned patches could try to make the fine-tuned model's activations look benign while still changing behavior, directly stressing the state-awareness mechanism; this attack class is absent from the 52-attack suite.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. AntiDote proposes a bi-level adversarial training procedure for open-weight LLMs. The defender is a LoRA adapter trained with a safety objective (a negative DPO loss evaluated under an adversarial patch) plus a decoupled capability objective computed on the clean, unattacked model. The adversary is a hypernetwork that maps the defender's internal activations to a LoRA patch selected to maximize the likelihood of harmful responses. The paper evaluates AntiDote on ten models ranging from 0.6B to 27B parameters, compares against SFT, RMU, Booster, TAR, RepNoise, and Vaccine, reports Harmful Score and Finetune Accuracy after 20:80 harmful-to-benign fine-tuning, presents a 52-attack red-teaming heatmap, and claims state-of-the-art robustness with less than 0.5% average utility degradation.

Significance. If the empirical claims were fully supported, the contribution would be valuable: AntiDote combines parameter-efficient training, a state-aware adversarial generator, and a decoupled capability loss, which is a plausible recipe for tamper resistance with a small utility cost. The evaluation is unusually broad for this literature, spanning ten architectures, six baselines, and 52 attack categories, and the two ablations directly target the two main design claims. However, the central threat-model claim is currently stronger than the evidence, and the safety evaluation is partly circular with the training data and classifier; the study is therefore best read as a demonstration of robustness against a specific tested attack distribution rather than against the unrestricted full-access adversary promised in Section 2.1.

major comments (3)
  1. [§2.1, §2.2, §4] The declared threat model in Eqs. (1)–(2) gives the adversary full parameter access and quantifies over all fine-tuning strategies A in A, but the method trains only against hypernetwork-generated rank-r LoRA patches (Eqs. (3)–(4)), and the main evaluation attacks in Section 4 are standard 20:80 harmful-to-benign fine-tuning runs rather than the hypernetwork adversary. No theorem or experiment shows that the hypernetwork patch family approximates the worst-case full fine-tuning, and no adaptive adversary targeting the defender LoRA, a full-parameter DPO attack maximizing Eq. (4), or a harmful-only (p=1.0) fine-tuning run is tested; Appendix E varies p only up to 0.2. The paper therefore establishes robustness against a specific attack distribution, not the unrestricted adversary promised in Section 2.1.
  2. [§3.1, §3.3] The safety training set D_safe is built from BeaverTails, the Harmful Score metric is computed by the BeaverTails classifier from (Ji et al. 2023), and the safety evaluation set includes BeaverTails items without any reported disjointness between train and test. Because the defense is trained on BeaverTails preference pairs and then scored by a BeaverTails-derived classifier, part of the HS improvement may reflect alignment with the training/evaluation distribution rather than transferable tamper-resistance. Reporting per-benchmark HS, a held-out BeaverTails split, or a classifier-agnostic metric would be needed to support the generalization claim.
  3. [Tables 4 and 5] The two ablation tables report conflicting Harmful Scores for the same 'Full' AntiDote method: Llama-3.2-3B is 5.7 in Table 4 but 8.5 in Table 5, Gemma-3-12B is 9.1 versus 5.9, and Gemma-3-27B is 9.9 versus 5.1. Since both tables are presented as evaluations of the full method and no differing experimental conditions are stated, the quantitative support for the state-awareness and decoupled-loss ablations is internally inconsistent and should be reconciled.
minor comments (5)
  1. [Fig. 2, Appendix C] The attack numbering is inconsistent: the main text and Figure 2 identify Adversarial Suffixes as Adv 19, while Appendix C lists Attack 19 as Base64 Encoding and Adversarial Suffix as Attack 26; please align the numbering so the heatmap columns are interpretable.
  2. [Abstract and Section 1] The abstract's 'up to 27.4% more robust' and the contribution's '78% reduction' are not derived from any explicit table calculation; please specify the reference baseline and the setting for each headline number.
  3. [Appendix E] Appendix E says 'As presented in Table 3' when it refers to the harmful-ratio table in the appendix, which is labeled Table 6; the cross-reference is wrong.
  4. [Conclusion] The conclusion states that 'a full discussion of limitations' is in the Appendix, but no limitations section appears there; please add it or remove the pointer.
  5. [§3.4] No code release or seed details are provided; adding them would materially aid reproducibility of the reported HS/FA numbers.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the defense objective, adversarial training, and external evaluation benchmarks are not definitionally tied to each other; the BeaverTails overlap is an evaluation nuance, not a by-construction reduction.

full rationale

The paper's derivation chain does not reduce to its inputs by construction. The adversary loss in Eq. (4) maximizes a DPO objective on the compromised model, while the defender loss in Eq. (5) minimizes the same objective on the same compromised model; this is a zero-sum training game, not a prediction that is definitionally equal to its training signal. The claimed robustness results are obtained after a separate, standard harmful fine-tuning attack (Section 4: "all models were fine-tuned on a dataset with a 20:80 mixture of harmful to benign data"), not by re-applying the hypernetwork's generated LoRA patch, so the reported Harmful Scores are not equal to the training objective by construction. The evaluation also includes external benchmarks such as StrongREJECT, HarmBench, XSTest, MMLU, HellaSwag, and GSM8K, which are independent of the training data and the BeaverTails classifier. The only overlap is that the safety training distribution includes BeaverTails and the HS metric uses the BeaverTails classifier from Ji et al. (2023), which can partially reflect alignment with the training distribution; however, this is an evaluation nuance rather than a by-construction equivalence because the HS score is computed on generated outputs after unseen fine-tuning and is supplemented by external benchmarks. The self-citation to Sanyal and Mandal (2025) appears only as background support for black-box guardrails and is not load-bearing for the central claim. No specific circular step meets the evidentiary threshold required by the review rules.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The framework relies on three categories of assumptions: proxy adequacy of the hypernetwork, coverage and reliability of the safety datasets and judge, and representativeness of the fixed attack in evaluation. These are reasonable starting points but they are not proven, and they set the ceiling on the tamper-resistance guarantee.

free parameters (3)
  • defender loss weight lambda = 0.8
    Weights the safety loss L_safe in L_defender; set empirically in Appendix ablations and not derived from theory.
  • KL regularization weight beta = 0.3
    Balances capability preservation against the base model; chosen by hand and validated in ablation.
  • LoRA rank = 16
    Determines capacity of both defender and adversary patches; chosen without explicit justification.
assumptions (4)
  • domain assumption The hypernetwork-generated LoRA patch is a sufficient proxy for the full fine-tuning adversary in Eq. 1.
    Section 2.2 replaces the intractable inner maximization with the hypernetwork; if this proxy is weak, the defender is not hardened against the true adversary.
  • domain assumption The BeaverTails and do-not-answer datasets span the relevant space of harmful behaviors.
    Section 3.1 uses 16 harm categories from these datasets for Dsafe; coverage determines generalization.
  • domain assumption The BeaverTails classifier reliably measures harmfulness for all 52 attacks.
    HS is computed with the Ji et al. classifier; any bias in the judge transfers to the reported robustness.
  • domain assumption The 20:80 harmful-to-benign fine-tuning mix is representative of realistic attacks.
    Section 4 uses this mix as the attack; stronger or adaptive attacks are not evaluated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs." pith.science (2026). https://pith.science/paper/XZMSQQEM

@misc{pith2026250908000,
  author       = {Pith},
  title        = {Pith review of: AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XZMSQQEM}},
  note         = {Machine review of arXiv:2509.08000}
}
read the original abstract

The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures struggle to preserve the general capabilities of the LLM while resisting a determined adversary with full access to the model's weights and architecture, who can use full-parameter fine-tuning to erase existing safeguards. To address this, we introduce AntiDote, a bi-level optimization procedure for training LLMs to be resistant to such tampering. AntiDote involves an auxiliary adversary hypernetwork that learns to generate malicious Low-Rank Adaptation (LoRA) weights conditioned on the defender model's internal activations. The defender LLM is then trained with an objective to nullify the effect of these adversarial weight additions, forcing it to maintain its safety alignment. We validate this approach against a diverse suite of 52 red-teaming attacks, including jailbreak prompting, latent space manipulation, and direct weight-space attacks. AntiDote is upto 27.4\% more robust against adversarial attacks compared to both tamper-resistance and unlearning baselines. Crucially, this robustness is achieved with a minimal trade-off in utility, incurring a performance degradation of upto less than 0.5\% across capability benchmarks including MMLU, HellaSwag, and GSM8K. Our work offers a practical and compute efficient methodology for building open-weight models where safety is a more integral and resilient property.

Figures

Figures reproduced from arXiv: 2509.08000 by the authors.

Figure 1
Figure 1. An overview of the Antidote training framework. When a harmful prompt is processed, the defender LLM’s internal activations are fed to an adversarial hypernetwork. The adversary is trained via DPO loss to generate a mali￾cious LoRA patch designed to compromise the defender. The defender, in turn, is trained with two distinct, decoupled ob￾jectives: a tamper-resistance loss computed on the attacked model to build res… view at source ↗
Figure 2
Figure 2. Per-Attack Harmfulness Comparison Across 52 Red-Teaming Attacks. A granular stress test evaluating each defense method against 52 distinct red-teaming vectors, from simple prompt injections to sophisticated, multi-layered attacks. A heatmap where rows are defenses and columns are attacks. Darker green indicates a lower, better Harmful Score (HS). The bottom bar for AntiDote demonstrates its broad-spectrum effectiven… view at source ↗
Figure 3
Figure 3. AntiDote Achieves State-of-the-Art Robustness Across Diverse Models. We compare the post-attack Harm￾ful Score (HS) of AntiDote against strong baselines after fine-tuning on a mixed dataset. While gradient-based meth￾ods like Booster effectively penalize the immediate harmful loss, AntiDote’s hypernetwork learns to recognize and coun￾teract the underlying compromised activation states that lead to failure. 4 Experim… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: AntiDote Breaks the Safety-Utility Trade-off Frontier. We evaluate the trade-off between model utility (Average FA) and safety (HS). Most defenses operate along a trade-off curve, sacrificing utility for safety. AntiDote places itself in the optimal quadrant because ou…
Figure 5
Figure 5. Figure 5: Qualitative example of robustness to obfuscation [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

91 extracted references · 20 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    AI, .; :; Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Wang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; Yu, K.; Liu, P.; Liu, Q.; Yue, S.; Yang, S.; Yang, S.; Xie, W.; Huang, W.; Hu, X.; Ren, X.; Niu, X.; Nie, P.; Li, Y.; Xu, Y.; Liu, Y.; Wang, Y.; Cai, Y.; Gu, Z.; Liu, Z.; and Dai, Z. 2025. Yi: Open Foundation Models by 01.AI. arXiv:2403.04652

  4. [4]

    Almazrouei, E.; Alobeidli, H.; Alshamsi, A.; Cappelli, A.; Cojocaru, R.; Debbah, M.; Étienne Goffinet; Hesslow, D.; Launay, J.; Malartic, Q.; Mazzotta, D.; Noune, B.; Pannier, B.; and Penedo, G. 2023. The Falcon Series of Open Language Models. arXiv:2311.16867

  5. [5]

    E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S

    Anil, C.; Durmus, E.; Sharma, M.; Benton, J.; Kundu, S.; Batson, J.; Rimsky, N.; Tong, M.; Mu, J.; Ford, D.; Mosconi, F.; Agrawal, R.; Schaeffer, R.; Bashkansky, N.; Svenningsen, S.; Lambert, M.; Radhakrishnan, A.; Denison, C. E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kapl...

  6. [6]

    Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Rimsky, N.; Gurnee, W.; and Nanda, N. 2024. Refusal in Language Models Is Mediated by a Single Direction. ArXiv, abs/2406.11717

  7. [7]

    Do Anything Now

    Backes, M.; Shen, Y.; Shen, X.; Zhang, Y.; and Chen, Z. J. 2023. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security

  8. [8]

    Bianchi, F.; Goldie, A.; Doumbouya, M. K. B.; Poesia, G.; Nandi, A.; Jurafsky, D.; Ghilardi, D.; and Manning, C. D. 2024. h4rm3l: A Dynamic Benchmark of Composable Jailbreak Attacks for LLM Safety Assessment. In International Conference on Learning Representations

Show all 91 references
  1. [9]

    A.; and Hill, E

    Buckmann, M.; Nguyen, Q. A.; and Hill, E. 2025. Revealing economic facts: LLMs know more than they say. ArXiv, abs/2505.08662

  2. [10]

    Carlini, N.; Liu, C.; Erlingsson, \'U .; Kos, J.; and Song, D. 2018. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks. In USENIX Security Symposium

  3. [11]

    Chang, W.; Zhu, T.; Zhao, Y.; Song, S.; Xiong, P.; Zhou, W.; and Li, Y. 2025. Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models. ArXiv, abs/2505.17519

  4. [12]

    J.; and Wong, E

    Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2023. Jailbreaking Black Box Large Language Models in Twenty Queries. 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 23--42

  5. [13]

    E.; Gandikota, R.; Ewart, A.; Rosati, D.; Wu, Z.; et al

    Che, Z.; Casper, S.; Kirk, R.; Satheesh, A.; Slocum, S.; McKinney, L. E.; Gandikota, R.; Ewart, A.; Rosati, D.; Wu, Z.; et al. 2025. Model tampering attacks enable more rigorous evaluations of llm capabilities. arXiv preprint arXiv:2502.05209

  6. [14]

    Chu, J.; Liu, Y.; Yang, Z.; Shen, X.; Backes, M.; and Zhang, Y. 2024. JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs. In Annual Meeting of the Association for Computational Linguistics

  7. [15]

    Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168

  8. [16]

    Cordonnier, J.-B.; Loukas, A.; and Jaggi, M. 2021. Multi-Head Attention: Collaborate Instead of Concatenate. arXiv:2006.16362

  9. [17]

    DeepSeek-AI; Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; Dai, D.; Guo, D.; Yang, D.; Chen, D.; Ji, D.; Li, E.; Lin, F.; Dai, F.; Luo, F.; Hao, G.; Chen, G.; Li, G.; Zhang, H.; Bao, H.; Xu, H.; Wang, H.; Zhang, H.; Ding, H.; Xi...

  10. [18]

    Ding, P.; Kuang, J.; Ma, D.; Cao, X.; Xian, Y.; Chen, J.; and Huang, S. 2023. A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. In North American Chapter of the Association for Computational Linguistics

  11. [19]

    C.; Allen, E

    Doerig, A.; Kietzmann, T. C.; Allen, E. J.; Wu, Y.; Naselaris, T.; Kay, K. N.; and Charest, I. 2022. Visual representations in the human brain are aligned with large language models

  12. [20]

    Dong, X.; Hu, W.; Xu, W.; and He, T. 2024. SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage. In Annual Meeting of the Association for Computational Linguistics

  13. [21]

    Dong, Y.; Shayegani, E.; and Abu-Ghazaleh, N. B. 2023. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In International Conference on Learning Representations

  14. [22]

    Ge, S.; Zhou, C.; Hou, R.; Khabsa, M.; Wang, Y.-C.; Wang, Q.; Han, J.; and Mao, Y. 2023. MART: Improving LLM Safety with Multi-round Automatic Red-Teaming. In North American Chapter of the Association for Computational Linguistics

  15. [23]

    Gong, Y.; Ran, D.; Liu, J.; Wang, C.; Cong, T.; Wang, A.; Duan, S.; and Wang, X. 2023. FigStep: Jailbreaking Large Vision-language Models via Typographic Visual Prompts. ArXiv, abs/2311.05608

  16. [24]

    Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; Yang, A.; Fan, A.; Goyal, A.; Hartshorn, A.; Yang, A.; Mitra, A.; Sravankumar, A.; Korenev, A.; Hinsvark, A.; Rao, A.; Zhang, A.; Rodriguez, A.; Gre...

  17. [25]

    Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security

  18. [26]

    K.; Lopez-Cuena, E.; Bayarri-Planas, J.; Tormos, A.; Hinjos, D.; Perez, P

    Gururajan, A. K.; Lopez-Cuena, E.; Bayarri-Planas, J.; Tormos, A.; Hinjos, D.; Perez, P. B.; Arias-Duart, A.; Martin-Torres, P. A.; Urcelay-Ganzabal, L.; Gonzalez-Mallo, M.; \'A lvarez-Napagao, S.; Ayguad'e-Parra, E.; and Garcia-Gasulla, U. C. D. 2024. Aloe: A Family of Fine-t...

  19. [27]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 a . Measuring Massive Multitask Language Understanding. arXiv:2009.03300

  20. [28]

    Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 b . Measuring Mathematical Problem Solving With the MATH Dataset. arXiv:2103.03874

  21. [29]

    Honovich, O.; Scialom, T.; Levy, O.; and Schick, T. 2022. Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor. arXiv:2212.09689

  22. [30]

    J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W

    Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685

  23. [32]

    F.; and Liu, L

    Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2025 b . Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation. arXiv:2409.01586

  24. [33]

    F.; and Liu, L

    Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2025 c . Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation. arXiv:2501.17433

  25. [35]

    Huang, T.; Hu, S.; and Liu, L. 2024 b . Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack. arXiv:2402.01109

  26. [36]

    Huang, Y.; Tang, J.; Chen, D.; Tang, B.; Wan, Y.; Sun, L.; and Zhang, X. 2024. Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings. In unknown

  27. [37]

    Jelodar, H.; Bai, S.; Hamedi, P.; Mohammadian, H.; Razavi-Far, R.; and Ghorbani, A. A. 2025. Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering. ArXiv, abs/2504.07137

  28. [38]

    Ji, J.; Liu, M.; Dai, J.; Pan, X.; Zhang, C.; Bian, C.; Zhang, C.; Sun, R.; Wang, Y.; and Yang, Y. 2023. BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset. arXiv:2307.04657

  29. [39]

    Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D

    Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023. Mistral 7B. arXiv:2310.06825

  30. [40]

    Jiang, F.; Xu, Z.; Niu, L.; Xiang, Z.; Ramasubramanian, B.; Li, B.; and Poovendran, R. 2024. ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs. In Annual Meeting of the Association for Computational Linguistics

  31. [41]

    Koloski, B.; Margeloiu, A.; Jiang, X.; Skrlj, B.; Simidjievski, N.; and Jamnik, M. 2025. LLM Embeddings for Deep Learning on Tabular Data. ArXiv, abs/2502.11596

  32. [42]

    Kuditipudi, R.; Thickstun, J.; Hashimoto, T.; and Liang, P. 2023. Robust distortion-free watermarks for language models. arXiv preprint arXiv:2307.15593

  33. [43]

    Lee, I.; and Seong, H. 2024. BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models. In arXiv.org

  34. [44]

    Li, H.; Guo, D.; Fan, W.; Xu, M.; Huang, J.; and Song, Y. 2023. Multi-step Jailbreaking Privacy Attacks on ChatGPT. ArXiv, abs/2304.05197

  35. [45]

    D.; and Dombrowski, A.-K

    Li, N.; Pan, A.; Gopal, A.; Yue, S.; Berrios, D.; Gatti, A.; Li, J. D.; and Dombrowski, A.-K. 2024. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv:2403.03218

  36. [47]

    Lin, X.; Acharya, M.; Roy, A.; and Jha, S. 2025. TeleLoRA: Teleporting Model-Specific Alignment Across LLMs. ArXiv, abs/2503.20228

  37. [48]

    Liu, A.; Tang, L.; Pan, T.; Yin, Y.; Wang, B.; and Yang, A. 2025. PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization. ArXiv, abs/2504.01444

  38. [49]

    Liu, T.; Li, X.; Yao, J.; Zhu, J.; Han, B.; and Zhou, Z. 2023 a . DeepInception: Hypnotize Large Language Model to Be Jailbreaker. ArXiv, abs/2311.03191

  39. [51]

    Liu, X.; Sun, T.; Xu, T.; Wu, F.; Wang, C.; Wang, X.; and Gao, J. 2024 b . SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation. ArXiv, abs/2406.12975

  40. [52]

    Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2023 b . AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. ArXiv, abs/2310.04451

  41. [55]

    Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; Forsyth, D.; and Hendrycks, D. 2024 c . HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. ArXiv, abs/2402.04249

  42. [56]

    R.; and Papernot, N

    Muresanu, A.; Thudi, A.; Zhang, M. R.; and Papernot, N. 2024. Unlearnable Algorithms for In-context Learning. arXiv:2402.00751

  43. [57]

    Nemecek, A.; Jiang, Y.; and Ayday, E. 2024. Topic-Based Watermarks for Large Language Models. arXiv preprint arXiv:2404.02138

  44. [58]

    J.; Hassani, H.; Robey, A.; and Wong, E

    Pappas, G. J.; Hassani, H.; Robey, A.; and Wong, E. 2023. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. Trans. Mach. Learn. Res., 2025

  45. [59]

    Pawelczyk, M.; Neel, S.; and Lakkaraju, H. 2024. In-Context Unlearning: Language Models as Few Shot Unlearners. arXiv:2310.07579

  46. [60]

    Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022. Red Teaming Language Models with Language Models. In Conference on Empirical Methods in Natural Language Processing

  47. [61]

    Primack, W.; Steneker, I.; Han, Z.; Goodside, R.; Yue, S.; Li, N.; Zhang, H.; Wang, Z.; and Menghini, C. 2024. LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet. ArXiv, abs/2408.15221

  48. [62]

    D.; and Finn, C

    Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2024. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290

  49. [63]

    Rao, A.; Vashistha, S.; Naik, A.; Aditya, S.; and Choudhury, M. 2023. Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks. In International Conference on Language Resources and Evaluation

  50. [64]

    R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D

    Röttger, P.; Kirk, H. R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. arXiv:2308.01263

  51. [65]

    Sanyal, D.; and Mandal, M. 2025. Agents Are All You Need for LLM Unlearning. In Second Conference on Language Modeling

  52. [66]

    Saparov, A.; and He, H. 2023. Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought. arXiv:2210.01240

  53. [67]

    Shah, R.; Feuillade-Montixi, Q.; Pour, S.; Tagade, A.; Casper, S.; and Rando, J. 2023. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation. ArXiv, abs/2311.03348

  54. [68]

    Shayegani, E.; Dong, Y.; and Abu-Ghazaleh, N. B. 2023. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In International Conference on Learning Representations

  55. [69]

    Shayegani, E.; Mamun, M. A. A.; Fu, Y.; Zaree, P.; Dong, Y.; and Abu-Ghazaleh, N. B. 2023. Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. ArXiv, abs/2310.10844

  56. [70]

    Do Anything Now

    Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2023. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security

  57. [71]

    Do Anything Now

    Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. arXiv:2308.03825

  58. [72]

    Shu, D.; Jin, M.; Chen, T.; Zhang, C.; and Zhang, Y. 2024. Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models. ArXiv, abs/2407.09292

  59. [73]

    Souly, A.; Lu, Q.; Bowen, D.; Trinh, T.; Hsieh, E.; Pandey, S.; Abbeel, P.; Svegliato, J.; Emmons, S.; Watkins, O.; and Toyer, S. 2024. A StrongREJECT for Empty Jailbreaks. arXiv:2402.10260

  60. [74]

    Tamirisa, R.; Bharathi, B.; Phan, L.; Zhou, A.; Gatti, A.; Suresh, T.; Lin, M.; Wang, J.; Wang, R.; Arel, R.; et al. 2024. Tamper-resistant safeguards for open-weight llms. arXiv preprint arXiv:2408.00761

  61. [75]

    M.; Goedeckemeyer, A.; Saade, A.; Feng, A.; Kolesnikov, A.; Bendebury, A.; Abdagic, A.; Vadi, A.; György, A.; Pinto, A

    Team, G.; Kamath, A.; Ferret, J.; Pathak, S.; Vieillard, N.; Merhej, R.; Perrin, S.; Matejovicova, T.; Ramé, A.; Rivière, M.; Rouillard, L.; Mesnard, T.; Cideron, G.; bastien Grill, J.; Ramos, S.; Yvinec, E.; Casbon, M.; Pot, E.; Penchev, I.; Liu, G.; Visin, F.; Kenealy, K.; B...

  62. [76]

    Verma, A.; Krishna, S.; Gehrmann, S.; Seshadri, M.; Pradhan, A.; Ault, T.; Barrett, L.; Rabinowitz, D.; Doucette, J.; and Phan, N. 2024. Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs). ArXiv, abs/2407.14937

  63. [77]

    M.; and Papadimitratos, P

    Wahr'eus, J.; Hussain, A. M.; and Papadimitratos, P. 2025. Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing. ArXiv, abs/2503.21598

  64. [78]

    Wang, Y.; Li, H.; Han, X.; Nakov, P.; and Baldwin, T. 2023. Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs. arXiv:2308.13387

  65. [80]

    Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023 b . Jailbroken: How Does LLM Safety Training Fail? ArXiv, abs/2307.02483

  66. [81]

    Wong, A.; Cao, H.; Liu, Z.; and Li, Y. 2024. SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis. ArXiv, abs/2410.15641

  67. [82]

    Wu, X.; Mao, X.; Li, F.; Zhang, X.; Li, X.; Teng, C.; Ji, D.; and Li, Z. 2025. TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis. In Annual Meeting of the Association for Computational Linguistics

  68. [83]

    Xiao, Z.; Held, W.; Liu, Y.; and Yang, D. 2023. Task-Agnostic Low-Rank Adapters for Unseen E nglish Dialects. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7857--7870. Singapore: Associatio...

  69. [84]

    Xu, H.; Zhang, W.; Wang, Z.; Xiao, F.; Zheng, R.; Feng, Y.; Ba, Z.; and Ren, K. 2024 a . RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent. ArXiv, abs/2407.16667

  70. [85]

    Xu, Z.; Liu, Y.; Deng, G.; Li, Y.; and Picek, S. 2024 b . A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. In Annual Meeting of the Association for Computational Linguistics

  71. [86]

    Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; ...

  72. [87]

    Yang, H.; Qu, L.; Shareghi, E.; and Haffari, G. 2024 a . Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models. ArXiv, abs/2410.11459

  73. [88]

    Yang, X.; Tang, X.; Han, J.; and Hu, S. 2024 b . The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models. ArXiv, abs/2411.11407

  74. [89]

    Yong, Z.-X.; Menghini, C.; and Bach, S. H. 2023. Low-Resource Languages Jailbreak GPT-4. ArXiv, abs/2310.02446

  75. [90]

    Youssef, P.; Zhao, Z.; Braun, D.; Schl \"o tterer, J.; and Seifert, C. 2025. Position: Editing large language models poses serious safety risks. arXiv preprint arXiv:2502.02958

  76. [92]

    Yu, Z.; Liu, X.; Liang, S.; Cameron, Z.; Xiao, C.; and Zhang, N. 2024 b . Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models. ArXiv, abs/2403.17336

  77. [93]

    Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? arXiv:1905.07830

  78. [95]

    Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; and Shi, W. 2024 b . How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs. ArXiv, abs/2401.06373

  79. [96]

    Zeng, Y.; Namkoong, H.; Lam, H.; and Liu, J. 2024 c . LLM Embeddings Improve Test-time Adaptation to Tabular Y|X-Shifts. ArXiv, abs/2410.07395

  80. [97]

    Zhang, R.; and Koushanfar, F. 2024. Watermarking Large Language Models and the Generated Content: Opportunities and Challenges. In 2024 58th Asilomar Conference on Signals, Systems, and Computers, 1779--1786. IEEE

  81. [98]

    Zhao, Z.; Meng, G.; Dong, Y.; Zhang, Y.; Liu, T.; and Chen, K. 2024. Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction. ArXiv, abs/2402.18104

  82. [99]

    Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023. LIMA: Less Is More for Alignment. arXiv:2305.11206

  83. [100]

    Üstün, A.; Aryabumi, V.; Yong, Z.-X.; Ko, W.-Y.; D'souza, D.; Onilude, G.; Bhandari, N.; Singh, S.; Ooi, H.-L.; Kayid, A.; Vargus, F.; Blunsom, P.; Longpre, S.; Muennighoff, N.; Fadaee, M.; Kreutzer, J.; and Hooker, S. 2024. Aya Model: An Instruction Finetuned Open-Access Mult...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.