REVIEW 3 major objections 5 minor 91 references
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read AntiDote uses a bi-level adversarial game to make open-weight LLMs resist malicious fine-tuning while keeping their skills.
desk verdict A genuinely new bi-level defense with well-tested design choices, but the evaluation does not actually measure the full-access worst-case adversary the threat model promises. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The adversarial hypernetwork is a multi-stage model (self-attention over layer activations, residual feed-forward blocks, and heterogeneous multi-headed LoRA output heads) that maps the defender's internal activation vectors to low-rank weight updates. It plays the role of a differentiable proxy for the inner-loop fine-tuning adversary, making the min-max objective in Equation 2 tractable and training the defender against a continuously adapting attack.
What would settle it
Fine-tune an AntiDote-hardened model on a harmful-benign fine-tuning mixture with a substantially higher harmful ratio or step budget than the paper's p=0.2 setting, with the attacker choosing updates that explicitly counter the defender's learned LoRA; if the Harmful Score returns to near the undefended SFT baseline, the proxy-fidelity assumption fails.
Extended reading notes
Core claim
AntiDote's central claim is that the intractable min-max problem of defending against an unrestricted fine-tuning adversary can be approximated by a fully differentiable bi-level game between an adversarial hypernetwork and a parameter-efficient defender. The hypernetwork consumes the defender's internal activations and emits malicious LoRA weights designed to make the attacked model prefer a harmful response; the defender's own LoRA weights are then trained to keep preferring the safe response even when that patch is applied. The paper further claims that computing the capability-preservation loss on the clean, unattacked model, decoupled from the safety loss, is what lets the method break the safety-utility trade-off. Empirically, the paper asserts that AntiDote achieves the lowest Harmful Score on all ten tested models while maintaining or improving fine-tune accuracy, and that its state-awareness is the load-bearing ingredient: replacing live activations with a static prompt embedding raises the Harmful Score three- to five-fold.
Load-bearing premise
The hypernetwork's low-rank weight patch is a faithful proxy for a determined full-parameter fine-tuning adversary; if a real attacker can do something the proxy does not capture, the claimed tamper-resistance could collapse.
Editorial extensions
If this is right
- AntiDote keeps safety alignment intact after a 20:80 harmful-benign fine-tuning mixture, reducing Harmful Score by up to 78% relative to SFT while matching or beating SFT fine-tune accuracy.
- The decoupled capability loss, computed on the clean model, is claimed to be what avoids the classic safety-utility trade-off; without it, fine-tune accuracy drops by more than five points on the largest tested model.
- The state-aware adversary, which attacks internal activations rather than prompt text, is claimed to be the reason robustness transfers across 52 red-teaming vectors, including attacks like role-playing and adversarial suffixes where gradient-based defenses are blind.
- Because both players are LoRA-based and the DPO reference state is reconstructed on the fly, the alignment stage requires modest compute, completing a 12B-model run in about the same wall-clock time as the Booster baseline while using less GPU memory.
Reading between the lines
- A logical extension the paper does not test is to scale the adversary to full-rank weight updates or condition it on activations from multiple layers; the paper's own state-awareness ablation suggests such a stronger adversary would force proportionally stronger defenses.
- The decoupled objective suggests a transferable recipe for other alignment interventions: separate conflicting objectives into distinct gradient streams computed on their own clean states, a recipe the paper does not test outside AntiDote.
- An adaptive attacker who knows the defender is trained against activation-conditioned patches could try to make the fine-tuned model's activations look benign while still changing behavior, directly stressing the state-awareness mechanism; this attack class is absent from the 52-attack suite.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. AntiDote proposes a bi-level adversarial training procedure for open-weight LLMs. The defender is a LoRA adapter trained with a safety objective (a negative DPO loss evaluated under an adversarial patch) plus a decoupled capability objective computed on the clean, unattacked model. The adversary is a hypernetwork that maps the defender's internal activations to a LoRA patch selected to maximize the likelihood of harmful responses. The paper evaluates AntiDote on ten models ranging from 0.6B to 27B parameters, compares against SFT, RMU, Booster, TAR, RepNoise, and Vaccine, reports Harmful Score and Finetune Accuracy after 20:80 harmful-to-benign fine-tuning, presents a 52-attack red-teaming heatmap, and claims state-of-the-art robustness with less than 0.5% average utility degradation.
Significance. If the empirical claims were fully supported, the contribution would be valuable: AntiDote combines parameter-efficient training, a state-aware adversarial generator, and a decoupled capability loss, which is a plausible recipe for tamper resistance with a small utility cost. The evaluation is unusually broad for this literature, spanning ten architectures, six baselines, and 52 attack categories, and the two ablations directly target the two main design claims. However, the central threat-model claim is currently stronger than the evidence, and the safety evaluation is partly circular with the training data and classifier; the study is therefore best read as a demonstration of robustness against a specific tested attack distribution rather than against the unrestricted full-access adversary promised in Section 2.1.
major comments (3)
- [§2.1, §2.2, §4] The declared threat model in Eqs. (1)–(2) gives the adversary full parameter access and quantifies over all fine-tuning strategies A in A, but the method trains only against hypernetwork-generated rank-r LoRA patches (Eqs. (3)–(4)), and the main evaluation attacks in Section 4 are standard 20:80 harmful-to-benign fine-tuning runs rather than the hypernetwork adversary. No theorem or experiment shows that the hypernetwork patch family approximates the worst-case full fine-tuning, and no adaptive adversary targeting the defender LoRA, a full-parameter DPO attack maximizing Eq. (4), or a harmful-only (p=1.0) fine-tuning run is tested; Appendix E varies p only up to 0.2. The paper therefore establishes robustness against a specific attack distribution, not the unrestricted adversary promised in Section 2.1.
- [§3.1, §3.3] The safety training set D_safe is built from BeaverTails, the Harmful Score metric is computed by the BeaverTails classifier from (Ji et al. 2023), and the safety evaluation set includes BeaverTails items without any reported disjointness between train and test. Because the defense is trained on BeaverTails preference pairs and then scored by a BeaverTails-derived classifier, part of the HS improvement may reflect alignment with the training/evaluation distribution rather than transferable tamper-resistance. Reporting per-benchmark HS, a held-out BeaverTails split, or a classifier-agnostic metric would be needed to support the generalization claim.
- [Tables 4 and 5] The two ablation tables report conflicting Harmful Scores for the same 'Full' AntiDote method: Llama-3.2-3B is 5.7 in Table 4 but 8.5 in Table 5, Gemma-3-12B is 9.1 versus 5.9, and Gemma-3-27B is 9.9 versus 5.1. Since both tables are presented as evaluations of the full method and no differing experimental conditions are stated, the quantitative support for the state-awareness and decoupled-loss ablations is internally inconsistent and should be reconciled.
minor comments (5)
- [Fig. 2, Appendix C] The attack numbering is inconsistent: the main text and Figure 2 identify Adversarial Suffixes as Adv 19, while Appendix C lists Attack 19 as Base64 Encoding and Adversarial Suffix as Attack 26; please align the numbering so the heatmap columns are interpretable.
- [Abstract and Section 1] The abstract's 'up to 27.4% more robust' and the contribution's '78% reduction' are not derived from any explicit table calculation; please specify the reference baseline and the setting for each headline number.
- [Appendix E] Appendix E says 'As presented in Table 3' when it refers to the harmful-ratio table in the appendix, which is labeled Table 6; the cross-reference is wrong.
- [Conclusion] The conclusion states that 'a full discussion of limitations' is in the Appendix, but no limitations section appears there; please add it or remove the pointer.
- [§3.4] No code release or seed details are provided; adding them would materially aid reproducibility of the reported HS/FA numbers.
Circularity Check
No significant circularity: the defense objective, adversarial training, and external evaluation benchmarks are not definitionally tied to each other; the BeaverTails overlap is an evaluation nuance, not a by-construction reduction.
full rationale
The paper's derivation chain does not reduce to its inputs by construction. The adversary loss in Eq. (4) maximizes a DPO objective on the compromised model, while the defender loss in Eq. (5) minimizes the same objective on the same compromised model; this is a zero-sum training game, not a prediction that is definitionally equal to its training signal. The claimed robustness results are obtained after a separate, standard harmful fine-tuning attack (Section 4: "all models were fine-tuned on a dataset with a 20:80 mixture of harmful to benign data"), not by re-applying the hypernetwork's generated LoRA patch, so the reported Harmful Scores are not equal to the training objective by construction. The evaluation also includes external benchmarks such as StrongREJECT, HarmBench, XSTest, MMLU, HellaSwag, and GSM8K, which are independent of the training data and the BeaverTails classifier. The only overlap is that the safety training distribution includes BeaverTails and the HS metric uses the BeaverTails classifier from Ji et al. (2023), which can partially reflect alignment with the training distribution; however, this is an evaluation nuance rather than a by-construction equivalence because the HS score is computed on generated outputs after unseen fine-tuning and is supplemented by external benchmarks. The self-citation to Sanyal and Mandal (2025) appears only as background support for black-box guardrails and is not load-bearing for the central claim. No specific circular step meets the evidentiary threshold required by the review rules.
Assumptions & free parameters
free parameters (3)
- defender loss weight lambda =
0.8
- KL regularization weight beta =
0.3
- LoRA rank =
16
assumptions (4)
- domain assumption The hypernetwork-generated LoRA patch is a sufficient proxy for the full fine-tuning adversary in Eq. 1.
- domain assumption The BeaverTails and do-not-answer datasets span the relevant space of harmful behaviors.
- domain assumption The BeaverTails classifier reliably measures harmfulness for all 52 attacks.
- domain assumption The 20:80 harmful-to-benign fine-tuning mix is representative of realistic attacks.
Cite this review
Pith. "Pith review of AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs." pith.science (2026). https://pith.science/paper/XZMSQQEM
@misc{pith2026250908000,
author = {Pith},
title = {Pith review of: AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/XZMSQQEM}},
note = {Machine review of arXiv:2509.08000}
}
read the original abstract
The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures struggle to preserve the general capabilities of the LLM while resisting a determined adversary with full access to the model's weights and architecture, who can use full-parameter fine-tuning to erase existing safeguards. To address this, we introduce AntiDote, a bi-level optimization procedure for training LLMs to be resistant to such tampering. AntiDote involves an auxiliary adversary hypernetwork that learns to generate malicious Low-Rank Adaptation (LoRA) weights conditioned on the defender model's internal activations. The defender LLM is then trained with an objective to nullify the effect of these adversarial weight additions, forcing it to maintain its safety alignment. We validate this approach against a diverse suite of 52 red-teaming attacks, including jailbreak prompting, latent space manipulation, and direct weight-space attacks. AntiDote is upto 27.4\% more robust against adversarial attacks compared to both tamper-resistance and unlearning baselines. Crucially, this robustness is achieved with a minimal trade-off in utility, incurring a performance degradation of upto less than 0.5\% across capability benchmarks including MMLU, HellaSwag, and GSM8K. Our work offers a practical and compute efficient methodology for building open-weight models where safety is a more integral and resilient property.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
AI, .; :; Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Wang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; Yu, K.; Liu, P.; Liu, Q.; Yue, S.; Yang, S.; Yang, S.; Xie, W.; Huang, W.; Hu, X.; Ren, X.; Niu, X.; Nie, P.; Li, Y.; Xu, Y.; Liu, Y.; Wang, Y.; Cai, Y.; Gu, Z.; Liu, Z.; and Dai, Z. 2025. Yi: Open Foundation Models by 01.AI. arXiv:2403.04652
arXiv 2025
-
[4]
Almazrouei, E.; Alobeidli, H.; Alshamsi, A.; Cappelli, A.; Cojocaru, R.; Debbah, M.; Étienne Goffinet; Hesslow, D.; Launay, J.; Malartic, Q.; Mazzotta, D.; Noune, B.; Pannier, B.; and Penedo, G. 2023. The Falcon Series of Open Language Models. arXiv:2311.16867
arXiv 2023
-
[5]
E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kaplan, J.; Ganguli, D.; Bowman, S
Anil, C.; Durmus, E.; Sharma, M.; Benton, J.; Kundu, S.; Batson, J.; Rimsky, N.; Tong, M.; Mu, J.; Ford, D.; Mosconi, F.; Agrawal, R.; Schaeffer, R.; Bashkansky, N.; Svenningsen, S.; Lambert, M.; Radhakrishnan, A.; Denison, C. E.; Hubinger, E.; Bai, Y.; Bricken, T.; Maxwell, T.; Schiefer, N.; Sully, J.; Tamkin, A.; Lanham, T.; Nguyen, K.; Korbak, T.; Kapl...
2024
-
[6]
Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Rimsky, N.; Gurnee, W.; and Nanda, N. 2024. Refusal in Language Models Is Mediated by a Single Direction. ArXiv, abs/2406.11717
arXiv 2024
-
[7]
Do Anything Now
Backes, M.; Shen, Y.; Shen, X.; Zhang, Y.; and Chen, Z. J. 2023. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security
2023
-
[8]
Bianchi, F.; Goldie, A.; Doumbouya, M. K. B.; Poesia, G.; Nandi, A.; Jurafsky, D.; Ghilardi, D.; and Manning, C. D. 2024. h4rm3l: A Dynamic Benchmark of Composable Jailbreak Attacks for LLM Safety Assessment. In International Conference on Learning Representations
2024
Show all 91 references
-
[9]
A.; and Hill, E
Buckmann, M.; Nguyen, Q. A.; and Hill, E. 2025. Revealing economic facts: LLMs know more than they say. ArXiv, abs/2505.08662
2025
-
[10]
Carlini, N.; Liu, C.; Erlingsson, \'U .; Kos, J.; and Song, D. 2018. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks. In USENIX Security Symposium
2018
-
[11]
Chang, W.; Zhu, T.; Zhao, Y.; Song, S.; Xiong, P.; Zhou, W.; and Li, Y. 2025. Chain-of-Lure: A Synthetic Narrative-Driven Approach to Compromise Large Language Models. ArXiv, abs/2505.17519
2025
-
[12]
J.; and Wong, E
Chao, P.; Robey, A.; Dobriban, E.; Hassani, H.; Pappas, G. J.; and Wong, E. 2023. Jailbreaking Black Box Large Language Models in Twenty Queries. 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 23--42
2023
-
[13]
E.; Gandikota, R.; Ewart, A.; Rosati, D.; Wu, Z.; et al
Che, Z.; Casper, S.; Kirk, R.; Satheesh, A.; Slocum, S.; McKinney, L. E.; Gandikota, R.; Ewart, A.; Rosati, D.; Wu, Z.; et al. 2025. Model tampering attacks enable more rigorous evaluations of llm capabilities. arXiv preprint arXiv:2502.05209
2025 arXiv
-
[14]
Chu, J.; Liu, Y.; Yang, Z.; Shen, X.; Backes, M.; and Zhang, Y. 2024. JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs. In Annual Meeting of the Association for Computational Linguistics
2024
-
[15]
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168
2021 arXiv
-
[16]
Cordonnier, J.-B.; Loukas, A.; and Jaggi, M. 2021. Multi-Head Attention: Collaborate Instead of Concatenate. arXiv:2006.16362
2021 arXiv
-
[17]
DeepSeek-AI; Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; Dai, D.; Guo, D.; Yang, D.; Chen, D.; Ji, D.; Li, E.; Lin, F.; Dai, F.; Luo, F.; Hao, G.; Chen, G.; Li, G.; Zhang, H.; Bao, H.; Xu, H.; Wang, H.; Zhang, H.; Ding, H.; Xi...
2025 arXiv
-
[18]
Ding, P.; Kuang, J.; Ma, D.; Cao, X.; Xian, Y.; Chen, J.; and Huang, S. 2023. A Wolf in Sheep’s Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. In North American Chapter of the Association for Computational Linguistics
2023
-
[19]
C.; Allen, E
Doerig, A.; Kietzmann, T. C.; Allen, E. J.; Wu, Y.; Naselaris, T.; Kay, K. N.; and Charest, I. 2022. Visual representations in the human brain are aligned with large language models
2022
-
[20]
Dong, X.; Hu, W.; Xu, W.; and He, T. 2024. SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage. In Annual Meeting of the Association for Computational Linguistics
2024
-
[21]
Dong, Y.; Shayegani, E.; and Abu-Ghazaleh, N. B. 2023. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In International Conference on Learning Representations
2023
-
[22]
Ge, S.; Zhou, C.; Hou, R.; Khabsa, M.; Wang, Y.-C.; Wang, Q.; Han, J.; and Mao, Y. 2023. MART: Improving LLM Safety with Multi-round Automatic Red-Teaming. In North American Chapter of the Association for Computational Linguistics
2023
-
[23]
Gong, Y.; Ran, D.; Liu, J.; Wang, C.; Cong, T.; Wang, A.; Duan, S.; and Wang, X. 2023. FigStep: Jailbreaking Large Vision-language Models via Typographic Visual Prompts. ArXiv, abs/2311.05608
2023 arXiv
-
[24]
Grattafiori, A.; Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Vaughan, A.; Yang, A.; Fan, A.; Goyal, A.; Hartshorn, A.; Yang, A.; Mitra, A.; Sravankumar, A.; Korenev, A.; Hinsvark, A.; Rao, A.; Zhang, A.; Rodriguez, A.; Gre...
2024 arXiv
-
[25]
Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; and Fritz, M. 2023. Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security
2023
-
[26]
K.; Lopez-Cuena, E.; Bayarri-Planas, J.; Tormos, A.; Hinjos, D.; Perez, P
Gururajan, A. K.; Lopez-Cuena, E.; Bayarri-Planas, J.; Tormos, A.; Hinjos, D.; Perez, P. B.; Arias-Duart, A.; Martin-Torres, P. A.; Urcelay-Ganzabal, L.; Gonzalez-Mallo, M.; \'A lvarez-Napagao, S.; Ayguad'e-Parra, E.; and Garcia-Gasulla, U. C. D. 2024. Aloe: A Family of Fine-t...
2024 arXiv
-
[27]
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 a . Measuring Massive Multitask Language Understanding. arXiv:2009.03300
2021 arXiv
-
[28]
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021 b . Measuring Mathematical Problem Solving With the MATH Dataset. arXiv:2103.03874
2021 arXiv
-
[29]
Honovich, O.; Scialom, T.; Levy, O.; and Schick, T. 2022. Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor. arXiv:2212.09689
2022 arXiv
-
[30]
J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685
2021 arXiv
-
[32]
F.; and Liu, L
Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2025 b . Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation. arXiv:2409.01586
2025 arXiv
-
[33]
F.; and Liu, L
Huang, T.; Hu, S.; Ilhan, F.; Tekin, S. F.; and Liu, L. 2025 c . Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation. arXiv:2501.17433
2025 arXiv
-
[35]
Huang, T.; Hu, S.; and Liu, L. 2024 b . Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack. arXiv:2402.01109
2024 arXiv
-
[36]
Huang, Y.; Tang, J.; Chen, D.; Tang, B.; Wan, Y.; Sun, L.; and Zhang, X. 2024. Jailbreaking Large Language Models Through Alignment Vulnerabilities in Out-of-Distribution Settings. In unknown
2024
-
[37]
Jelodar, H.; Bai, S.; Hamedi, P.; Mohammadian, H.; Razavi-Far, R.; and Ghorbani, A. A. 2025. Large Language Model (LLM) for Software Security: Code Analysis, Malware Analysis, Reverse Engineering. ArXiv, abs/2504.07137
2025 arXiv
-
[38]
Ji, J.; Liu, M.; Dai, J.; Pan, X.; Zhang, C.; Bian, C.; Zhang, C.; Sun, R.; Wang, Y.; and Yang, Y. 2023. BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset. arXiv:2307.04657
2023
-
[39]
Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; de las Casas, D.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; Lavaud, L. R.; Lachaux, M.-A.; Stock, P.; Scao, T. L.; Lavril, T.; Wang, T.; Lacroix, T.; and Sayed, W. E. 2023. Mistral 7B. arXiv:2310.06825
2023 arXiv
-
[40]
Jiang, F.; Xu, Z.; Niu, L.; Xiang, Z.; Ramasubramanian, B.; Li, B.; and Poovendran, R. 2024. ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs. In Annual Meeting of the Association for Computational Linguistics
2024
-
[41]
Koloski, B.; Margeloiu, A.; Jiang, X.; Skrlj, B.; Simidjievski, N.; and Jamnik, M. 2025. LLM Embeddings for Deep Learning on Tabular Data. ArXiv, abs/2502.11596
2025 arXiv
-
[42]
Kuditipudi, R.; Thickstun, J.; Hashimoto, T.; and Liang, P. 2023. Robust distortion-free watermarks for language models. arXiv preprint arXiv:2307.15593
2023 arXiv
-
[43]
Lee, I.; and Seong, H. 2024. BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models. In arXiv.org
2024
-
[44]
Li, H.; Guo, D.; Fan, W.; Xu, M.; Huang, J.; and Song, Y. 2023. Multi-step Jailbreaking Privacy Attacks on ChatGPT. ArXiv, abs/2304.05197
2023 arXiv
-
[45]
D.; and Dombrowski, A.-K
Li, N.; Pan, A.; Gopal, A.; Yue, S.; Berrios, D.; Gatti, A.; Li, J. D.; and Dombrowski, A.-K. 2024. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv:2403.03218
2024 arXiv
-
[47]
Lin, X.; Acharya, M.; Roy, A.; and Jha, S. 2025. TeleLoRA: Teleporting Model-Specific Alignment Across LLMs. ArXiv, abs/2503.20228
2025 arXiv
-
[48]
Liu, A.; Tang, L.; Pan, T.; Yin, Y.; Wang, B.; and Yang, A. 2025. PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization. ArXiv, abs/2504.01444
2025
-
[49]
Liu, T.; Li, X.; Yao, J.; Zhu, J.; Han, B.; and Zhou, Z. 2023 a . DeepInception: Hypnotize Large Language Model to Be Jailbreaker. ArXiv, abs/2311.03191
2023 arXiv
-
[51]
Liu, X.; Sun, T.; Xu, T.; Wu, F.; Wang, C.; Wang, X.; and Gao, J. 2024 b . SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation. ArXiv, abs/2406.12975
2024 arXiv
-
[52]
Liu, X.; Xu, N.; Chen, M.; and Xiao, C. 2023 b . AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. ArXiv, abs/2310.04451
2023 arXiv
-
[55]
Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; Forsyth, D.; and Hendrycks, D. 2024 c . HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. ArXiv, abs/2402.04249
2024 arXiv
-
[56]
R.; and Papernot, N
Muresanu, A.; Thudi, A.; Zhang, M. R.; and Papernot, N. 2024. Unlearnable Algorithms for In-context Learning. arXiv:2402.00751
2024
-
[57]
Nemecek, A.; Jiang, Y.; and Ayday, E. 2024. Topic-Based Watermarks for Large Language Models. arXiv preprint arXiv:2404.02138
2024 arXiv
-
[58]
J.; Hassani, H.; Robey, A.; and Wong, E
Pappas, G. J.; Hassani, H.; Robey, A.; and Wong, E. 2023. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. Trans. Mach. Learn. Res., 2025
2023
-
[59]
Pawelczyk, M.; Neel, S.; and Lakkaraju, H. 2024. In-Context Unlearning: Language Models as Few Shot Unlearners. arXiv:2310.07579
2024 arXiv
-
[60]
Perez, E.; Huang, S.; Song, F.; Cai, T.; Ring, R.; Aslanides, J.; Glaese, A.; McAleese, N.; and Irving, G. 2022. Red Teaming Language Models with Language Models. In Conference on Empirical Methods in Natural Language Processing
2022
-
[61]
Primack, W.; Steneker, I.; Han, Z.; Goodside, R.; Yue, S.; Li, N.; Zhang, H.; Wang, Z.; and Menghini, C. 2024. LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet. ArXiv, abs/2408.15221
2024 arXiv
-
[62]
D.; and Finn, C
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2024. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arXiv:2305.18290
2024 arXiv
-
[63]
Rao, A.; Vashistha, S.; Naik, A.; Aditya, S.; and Choudhury, M. 2023. Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks. In International Conference on Language Resources and Evaluation
2023
-
[64]
R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D
Röttger, P.; Kirk, H. R.; Vidgen, B.; Attanasio, G.; Bianchi, F.; and Hovy, D. 2024. XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models. arXiv:2308.01263
2024 arXiv
-
[65]
Sanyal, D.; and Mandal, M. 2025. Agents Are All You Need for LLM Unlearning. In Second Conference on Language Modeling
2025
-
[66]
Saparov, A.; and He, H. 2023. Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought. arXiv:2210.01240
2023 arXiv
-
[67]
Shah, R.; Feuillade-Montixi, Q.; Pour, S.; Tagade, A.; Casper, S.; and Rando, J. 2023. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation. ArXiv, abs/2311.03348
2023 arXiv
-
[68]
Shayegani, E.; Dong, Y.; and Abu-Ghazaleh, N. B. 2023. Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models. In International Conference on Learning Representations
2023
-
[69]
Shayegani, E.; Mamun, M. A. A.; Fu, Y.; Zaree, P.; Dong, Y.; and Abu-Ghazaleh, N. B. 2023. Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks. ArXiv, abs/2310.10844
2023 arXiv
-
[70]
Do Anything Now
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2023. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security
2023
-
[71]
Do Anything Now
Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; and Zhang, Y. 2024. "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. arXiv:2308.03825
2024 arXiv
-
[72]
Shu, D.; Jin, M.; Chen, T.; Zhang, C.; and Zhang, Y. 2024. Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models. ArXiv, abs/2407.09292
2024 arXiv
-
[73]
Souly, A.; Lu, Q.; Bowen, D.; Trinh, T.; Hsieh, E.; Pandey, S.; Abbeel, P.; Svegliato, J.; Emmons, S.; Watkins, O.; and Toyer, S. 2024. A StrongREJECT for Empty Jailbreaks. arXiv:2402.10260
2024 arXiv
-
[74]
Tamirisa, R.; Bharathi, B.; Phan, L.; Zhou, A.; Gatti, A.; Suresh, T.; Lin, M.; Wang, J.; Wang, R.; Arel, R.; et al. 2024. Tamper-resistant safeguards for open-weight llms. arXiv preprint arXiv:2408.00761
2024 arXiv
-
[75]
M.; Goedeckemeyer, A.; Saade, A.; Feng, A.; Kolesnikov, A.; Bendebury, A.; Abdagic, A.; Vadi, A.; György, A.; Pinto, A
Team, G.; Kamath, A.; Ferret, J.; Pathak, S.; Vieillard, N.; Merhej, R.; Perrin, S.; Matejovicova, T.; Ramé, A.; Rivière, M.; Rouillard, L.; Mesnard, T.; Cideron, G.; bastien Grill, J.; Ramos, S.; Yvinec, E.; Casbon, M.; Pot, E.; Penchev, I.; Liu, G.; Visin, F.; Kenealy, K.; B...
2025 arXiv
-
[76]
Verma, A.; Krishna, S.; Gehrmann, S.; Seshadri, M.; Pradhan, A.; Ault, T.; Barrett, L.; Rabinowitz, D.; Doucette, J.; and Phan, N. 2024. Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs). ArXiv, abs/2407.14937
2024
-
[77]
M.; and Papadimitratos, P
Wahr'eus, J.; Hussain, A. M.; and Papadimitratos, P. 2025. Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing. ArXiv, abs/2503.21598
2025 arXiv
-
[78]
Wang, Y.; Li, H.; Han, X.; Nakov, P.; and Baldwin, T. 2023. Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs. arXiv:2308.13387
2023 arXiv
-
[80]
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2023 b . Jailbroken: How Does LLM Safety Training Fail? ArXiv, abs/2307.02483
2023 arXiv
-
[81]
Wong, A.; Cao, H.; Liu, Z.; and Li, Y. 2024. SMILES-Prompting: A Novel Approach to LLM Jailbreak Attacks in Chemical Synthesis. ArXiv, abs/2410.15641
2024 arXiv
-
[82]
Wu, X.; Mao, X.; Li, F.; Zhang, X.; Li, X.; Teng, C.; Ji, D.; and Li, Z. 2025. TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis. In Annual Meeting of the Association for Computational Linguistics
2025
-
[83]
Xiao, Z.; Held, W.; Liu, Y.; and Yang, D. 2023. Task-Agnostic Low-Rank Adapters for Unseen E nglish Dialects. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7857--7870. Singapore: Associatio...
2023
-
[84]
Xu, H.; Zhang, W.; Wang, Z.; Xiao, F.; Zheng, R.; Feng, Y.; Ba, Z.; and Ren, K. 2024 a . RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent. ArXiv, abs/2407.16667
2024 arXiv
-
[85]
Xu, Z.; Liu, Y.; Deng, G.; Li, Y.; and Picek, S. 2024 b . A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. In Annual Meeting of the Association for Computational Linguistics
2024
-
[86]
Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; ...
2025 arXiv
-
[87]
Yang, H.; Qu, L.; Shareghi, E.; and Haffari, G. 2024 a . Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models. ArXiv, abs/2410.11459
2024 arXiv
-
[88]
Yang, X.; Tang, X.; Han, J.; and Hu, S. 2024 b . The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models. ArXiv, abs/2411.11407
2024 arXiv
-
[89]
Yong, Z.-X.; Menghini, C.; and Bach, S. H. 2023. Low-Resource Languages Jailbreak GPT-4. ArXiv, abs/2310.02446
2023 arXiv
-
[90]
Youssef, P.; Zhao, Z.; Braun, D.; Schl \"o tterer, J.; and Seifert, C. 2025. Position: Editing large language models poses serious safety risks. arXiv preprint arXiv:2502.02958
2025 arXiv
-
[92]
Yu, Z.; Liu, X.; Liang, S.; Cameron, Z.; Xiao, C.; and Zhang, N. 2024 b . Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models. ArXiv, abs/2403.17336
2024 arXiv
-
[93]
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? arXiv:1905.07830
2019 arXiv
-
[95]
Zeng, Y.; Lin, H.; Zhang, J.; Yang, D.; Jia, R.; and Shi, W. 2024 b . How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs. ArXiv, abs/2401.06373
2024 arXiv
-
[96]
Zeng, Y.; Namkoong, H.; Lam, H.; and Liu, J. 2024 c . LLM Embeddings Improve Test-time Adaptation to Tabular Y|X-Shifts. ArXiv, abs/2410.07395
2024 arXiv
-
[97]
Zhang, R.; and Koushanfar, F. 2024. Watermarking Large Language Models and the Generated Content: Opportunities and Challenges. In 2024 58th Asilomar Conference on Signals, Systems, and Computers, 1779--1786. IEEE
2024
-
[98]
Zhao, Z.; Meng, G.; Dong, Y.; Zhang, Y.; Liu, T.; and Chen, K. 2024. Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction. ArXiv, abs/2402.18104
2024 arXiv
-
[99]
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023. LIMA: Less Is More for Alignment. arXiv:2305.11206
2023 arXiv
-
[100]
Üstün, A.; Aryabumi, V.; Yong, Z.-X.; Ko, W.-Y.; D'souza, D.; Onilude, G.; Bhandari, N.; Singh, S.; Ooi, H.-L.; Kayid, A.; Vargus, F.; Blunsom, P.; Longpre, S.; Muennighoff, N.; Fadaee, M.; Kreutzer, J.; and Hooker, S. 2024. Aya Model: An Instruction Finetuned Open-Access Mult...
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.