REVIEW 2 major objections 6 minor 71 references
Jailbreaks suppress either the harmfulness or refusal direction inside LLMs before any token is generated; coupling those directions at both prompt and response positions hardens safety without the usual capability tax.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
HARC couples harmfulness and refusal directions at prompt and response positions, yielding the best robustness-capability-usability trade-off among major safety methods.
T0 review reviewed 2026-07-12 challenge →
load-bearing objection Solid mechanistic paper that actually ships a usable dual-position coupling method with the best reported robustness/capability/over-refusal trade-off among the usual baselines; residual CodeAttack and FT-attack fragility are real but already scoped by the authors. the 2 major comments →
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Successful jailbreaks suppress the refusal direction, the harmfulness direction, or both during prompt encoding, locking the model onto a compliance trajectory before generation starts; yet separable harmfulness and refusal signals reappear at response-token positions even when prompt-side recognition failed. Pairing those four directions (prompt/response × harm/refusal) with a subspace-confined coupling loss restores refusal without degrading general capability or inflating over-refusal, producing the strongest robustness–capability–usability trade-off among six major training- and inference-time safety baselines.
What carries the argument
HARC (Harmfulness-And-Refusal Coupling): LoRA fine-tuning that applies additive-margin hinge losses on cosine projections of residual activations onto the harmfulness and refusal directions at both the last-instruction / post-instruction tokens and mean-pooled early response tokens, plus a heavy KL term on benign data and cross-entropy on refusal text. The intervention is deliberately restricted to the two-dimensional harmfulness–refusal subspace; directions are periodically recomputed and EMA-blended as the residual geometry shifts.
Load-bearing premise
The coupling loss can only amplify an already-present projection onto the extracted harmfulness direction; if an attack drives residuals nearly orthogonal to both directions at both prompt and response positions, the gradient vanishes and the defense has almost nothing to pull on.
What would settle it
After HARC training, measure attack success on a held-out obfuscation or rewrite family whose residuals stay near zero (or negative) on both harmfulness and refusal across layers; if that family’s ASR remains high while DAN/PAIR/DeepInception collapse, dual-position coupling is insufficient for attacks outside the span of the extracted directions.
If this is right
- Black-box jailbreaks can be diagnosed by which region of the harmfulness–refusal plane they occupy at prompt encoding.
- Safety methods that act only at the prompt boundary miss the response-side recognition signal that often still fires during compliant generation.
- Confining fine-tuning to a two-dimensional safety subspace can avoid the alignment tax that broader residual-stream interventions incur.
- The four-direction structure is shared enough across model families that a single layer-selection score transfers without per-architecture retuning.
- Attacks that keep residuals orthogonal to both extracted directions (e.g., code obfuscation) remain the residual hard case for coupling-style defenses.
Where Pith is reading between the lines
- Direction extraction sets should deliberately include obfuscation and rewrite distributions so those patterns lie inside the span of v_harm rather than orthogonal to it.
- An adaptive attacker who knows the coupling target could deliberately keep residuals off the trained subspace, creating a cat-and-mouse dynamic analogous to adversarial training.
- The same dual-position coupling idea may apply to other dissociated concept pairs (truthfulness vs. sycophancy, instruction-following vs. flattery) if they show analogous prompt/response geometry.
- Because HARC’s footprint is low-dimensional, weight-access adversaries can undo it with few harmful examples; broader residual rerouting may be preferable when model weights are in the threat model.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that aligned LLMs encode harmfulness and refusal as separable residual-stream directions at both prompt-side and response-side token positions, that successful jailbreaks suppress one or both prompt-side directions before generation, and that the model can still recognize harm while generating even when prompt-side refusal failed. Motivated by this four-direction geometry, the authors introduce HARC, a LoRA fine-tuning method that couples the two directions at both positions via an additive-margin hinge loss on cosine projections, confined to the harmfulness–refusal subspace and combined with KL retention and refusal CE. Across Llama-3.1-8B/70B and Qwen-2.5-7B/72B, with transfer checks on three further families, HARC is reported to cut mean jailbreak ASR by roughly 4.7× relative to the base model while matching or undercutting base over-refusal and preserving capability, yielding the best robustness–capability–usability trade-off among six training- and inference-time baselines (Table 1, Table 7).
Significance. If the results hold, the paper makes two contributions of clear value to AI safety and mechanistic interpretability: (i) a multi-architecture characterization that jailbreaks occupy separable regions of a harmfulness–refusal plane and that response-side harm recognition can persist after prompt-side failure, and (ii) a subspace-targeted training intervention that improves black-box jailbreak robustness without the usual capability or over-refusal tax. Strengths include a transferable layer-selection score (Eq. 6), dual-position ablations (Table 6), multi-seed variance (Appendix E), five-family geometry transfer (Appendix A.2), 70B/72B scaling (Figure 5, Table 7), an explicit threat model, and released code. The work is therefore both diagnostically informative and practically actionable under the stated prompt-space threat model.
major comments (2)
- §5.2 and Table 6: the mechanistic motivation for dual-position coupling is that some attacks suppress prompt-side harm entirely and only fire at response positions (§3.2, §4). The ablation, however, shows that prompt-only and response-only variants already achieve harmfulness means comparable to Full HARC on Llama (0.120 / 0.113 vs 0.120) and that the dual objective’s distinctive win is over-refusal (XSTest 0.035). Please either (a) strengthen the claim that dual-position is required for robustness on attack classes that only activate response-side directions (e.g., report per-attack breakdowns for the single-position variants, especially CodeAttack), or (b) reframe the dual-position design as primarily a usability/over-refusal regularizer rather than a necessary condition for catching response-side-only harm.
- §4.2, Limitations §7, and Table 1: the paper correctly notes that the coupling gradient vanishes when residuals are near-orthogonal to both extracted directions, which is why CodeAttack remains the residual hard case (Llama 0.350→0.290; Qwen 0.417→0.340). Given that the abstract and §5.2 headline “strongest robustness… trade-off,” the main text should more explicitly bound the claim to attack distributions that already project onto the extracted directions, and state that direction-extraction coverage (including obfuscation-style attacks) is part of the method’s effectiveness, not an optional future extension. Without that qualification, the robustness claim overstates what the loss can do under the paper’s own geometry.
minor comments (6)
- Figure 2 caption and §3.2: “peak token” selection for response-side projections is deferred to Appendix A.3; a one-sentence definition in the main text would help readers interpret the scatter plots without flipping.
- Eq. (2): clarify that harmful continuations are obtained by ablating v_ref (as stated later in Appendix B) at the first mention, so the construction of v_resp_harm / v_resp_ref is self-contained.
- Table 1: “Ours + DPO” is the strongest Llama configuration on ASR; the abstract’s “HARC achieves the strongest…” phrasing should note when the hybrid is the reported best, or restrict the claim to the pure coupling method.
- §5.1 / Appendix D: free parameters (m, λ’s, β, K ramp) are only partially ablated (loss components in Table 6). A short sensitivity note for m and λ_kl would strengthen reproducibility claims.
- Appendix I’s fine-tuning-attack fragility is important for deployment readers; a brief pointer in the main Limitations paragraph would improve balance without expanding the threat model.
- Typos / polish: “arXiv:2607.00572v3” date line; occasional missing spaces before citations; ensure consistent notation for v_resp_harm vs v^{resp}_harm across figures and equations.
Circularity Check
No circularity: HARC is an empirical subspace fine-tuning method evaluated on independent jailbreak and capability benchmarks; nothing reduces by construction.
full rationale
The paper's chain is observational then interventional, not definitional. Harmfulness/refusal directions are extracted by difference-of-means (Eq. 1–2) on held-out AdvBench/UltraChat prompts; the coupling objective (Eq. 3–5) then reshapes LoRA residuals toward those (periodically EMA-refreshed) fixed vectors. Reported gains—ASR on JailbreakBench attacks (PAIR, PAP, DeepInception, CodeAttack), over-refusal on XSTest/CoCoNot, and capability on MMLU/GSM8K/HumanEval/IFEval/MT-Bench—are external LLM-judge and accuracy metrics, not algebraic rearrangements of the extraction sets or loss weights. Layer selection (Eq. 6) is a geometric heuristic for where to apply the loss, not a prediction forced by the fit. Prior citations (Arditi et al., Zhao et al.) are external authors; no self-citation uniqueness theorem or ansatz is load-bearing. Ablations (Table 6), multi-seed runs, five-family transfer, and 70B/72B scaling further treat outcomes as empirical. CodeAttack residual failure is acknowledged as a limitation of gradient signal, not hidden circularity. Score 0 is appropriate.
Axiom & Free-Parameter Ledger
free parameters (4)
- coupling margin m =
0.5
- loss weights λ_c, λ_cr, λ_kl, λ_ce =
λ_kl=10, others=1
- EMA β and recompute interval =
β=0.3, every 200 steps
- top-K layer ramp (2→4) =
K=2→4
axioms (3)
- domain assumption High-level concepts (harmfulness, refusal) are linearly represented as directions in residual-stream activations (difference-of-means).
- ad hoc to paper Mean-pooling the first 32 response tokens yields a stable response-side direction.
- domain assumption Confining the intervention to the two-dimensional harm–refusal subspace leaves the rest of the residual stream (and therefore capability) essentially undisturbed.
invented entities (2)
-
response-side harmfulness and refusal directions (v_resp_harm, v_resp_ref)
independent evidence
-
HARC coupling objective (additive-margin hinge on cosine projections at both positions)
no independent evidence
Cite this review
Pith. "Pith review of HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment." pith.science (2026). https://pith.science/paper/2BBQWXPK
@misc{pith2026260700572,
author = {Pith},
title = {Pith review of: HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/2BBQWXPK}},
note = {Machine review of arXiv:2607.00572}
}
read the original abstract
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.
Figures
Reference graph
Works this paper leans on
-
[1]
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043, 2023
Pith/arXiv arXiv 2023
-
[2]
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Autodan: Generating stealthy jailbreak prompts on aligned large language models. InInt. Conf. Learn. Rep. (ICLR), 2023
2023
-
[3]
How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 14322–14350, 2024
2024
-
[4]
do anything now
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. InProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 1671–1685, 2024
2024
-
[5]
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 23–42. IEEE, 2025
2025
-
[6]
Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms
Xiaogeng Liu, Peiran Li, Edward Suh, Yevgeniy V orobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao. Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. InInt. Conf. Learn. Rep. (ICLR), pages 22337–22384, 2025
2025
-
[7]
Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack
Mark Russinovich, Ahmed Salem, and Ronen Eldan. Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack. In34th USENIX Security Symposium (USENIX Security 25), pages 2421–2440, 2025
2025
-
[8]
Between a rock and a hard place: The tension between ethical reasoning and safety alignment in llms
Shei Pern Chua, Zhen Leng Thai, Kai Jun Teh, Xiao Li, Qibing Ren, and Xiaolin Hu. Between a rock and a hard place: The tension between ethical reasoning and safety alignment in llms. arXiv preprint arXiv:2509.05367, 2025
Pith/arXiv arXiv 2025
-
[9]
Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. Deepinception: Hypnotize large language model to be jailbreaker.arXiv preprint arXiv:2311.03191, 2023
Pith/arXiv arXiv 2023
-
[10]
Codeat- tack: Revealing safety generalization challenges of large language models via code completion
Qibing Ren, Chang Gao, Jing Shao, Junchi Yan, Xin Tan, Wai Lam, and Lizhuang Ma. Codeat- tack: Revealing safety generalization challenges of large language models via code completion. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 11437–11452, 2024
2024
-
[11]
Artprompt: Ascii art-based jailbreak attacks against aligned llms
Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran. Artprompt: Ascii art-based jailbreak attacks against aligned llms. In Ann. Meet. Assoc. Comput. Linguistics (ACL), pages 15157–15173, 2024
2024
-
[12]
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. InInt. Conf. Mach. Learn. (ICML), 2024
2024
-
[13]
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Adv. Neural Inform. Process. Syst. (NeurIPS), 36:53728–53741, 2023
2023
-
[14]
Safety alignment should be made more than just a few tokens deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson. Safety alignment should be made more than just a few tokens deep. InInt. Conf. Learn. Rep. (ICLR), 2024
2024
-
[15]
Melody Y Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, et al. Deliberative alignment: Reasoning enables safer language models.arXiv preprint arXiv:2412.16339, 2024
Pith/arXiv arXiv 2024
-
[16]
Stair: Improving safety alignment with introspective reasoning
Yichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia, Zhengwei Fang, Xiao Yang, Ranjie Duan, Dong Yan, Yinpeng Dong, and Jun Zhu. Stair: Improving safety alignment with introspective reasoning. InInt. Conf. Mach. Learn. (ICML), pages 76754–76777, 2025. 11
2025
-
[17]
Programming refusal with conditional activation steering
Bruce W Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Erik Miehling, Pierre Dognin, Manish Nagireddy, and Amit Dhurandhar. Programming refusal with conditional activation steering. InInt. Conf. Learn. Rep. (ICLR), 2025
2025
-
[18]
Scans: Mitigating the exaggerated safety for llms via safety-conscious activation steering
Zouying Cao, Yifei Yang, and Hai Zhao. Scans: Mitigating the exaggerated safety for llms via safety-conscious activation steering. InAAAI, volume 39, pages 23523–23531, 2025
2025
-
[19]
Improving alignment and robustness with circuit breakers.Adv
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks. Improving alignment and robustness with circuit breakers.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:83345–83373, 2024
2024
-
[20]
Representation bending for large language model safety
Ashkan Yousefpour, Taeheon Kim, Ryan Sungmo Kwon, Seungbeen Lee, Wonje Jeung, Seungju Han, Alvin Wan, Harrison Ngan, Youngjae Yu, and Jonghyun Choi. Representation bending for large language model safety. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 24073–24098, 2025
2025
-
[21]
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al. Latent adver- sarial training improves robustness to persistent harmful behaviors in llms.arXiv preprint arXiv:2407.15549, 2024
Pith/arXiv arXiv 2024
-
[22]
On effects of steering latent representation for large language model unlearning
Huu-Tien Dang, Tin Pham, Hoang Thanh-Tung, and Naoya Inoue. On effects of steering latent representation for large language model unlearning. InAAAI, volume 39, pages 23733–23742, 2025
2025
-
[23]
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction. InAdv. Neural Inform. Process. Syst. (NeurIPS), 2024
2024
-
[24]
Llms encode harmful- ness and refusal separately
Jiachen Zhao, Jing Huang, Zhengxuan Wu, David Bau, and Weiyan Shi. Llms encode harmful- ness and refusal separately. InAdv. Neural Inform. Process. Syst. (NeurIPS), 2025
2025
-
[25]
A general language assistant as a laboratory for alignment.arXiv preprint arXiv:2112.00861, 2021
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al. A general language assistant as a laboratory for alignment.arXiv preprint arXiv:2112.00861, 2021
Pith/arXiv arXiv 2021
-
[26]
Training language models to follow instructions with human feedback.Adv
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Adv. Neural Inform. Process. Syst. (NeurIPS), 35: 27730–27744, 2022
2022
-
[27]
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback.arXiv preprint arXiv:2204.05862, 2022
Pith/arXiv arXiv 2022
-
[28]
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models. InInt. Conf. Mach. Learn. (ICML), 2024
2024
-
[29]
Toy models of superposition.arXiv preprint arXiv:2209.10652, 2022
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al. Toy models of superposition.arXiv preprint arXiv:2209.10652, 2022
Pith/arXiv arXiv 2022
-
[30]
Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv preprint arXiv:2310.06824, 2023
Pith/arXiv arXiv 2023
-
[31]
Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering.arXiv preprint arXiv:2308.10248, 2023
Pith/arXiv arXiv 2023
-
[32]
Steering llama 2 via contrastive activation addition
Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 15504–15522, 2024. 12
2024
-
[33]
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency.arXiv preprint arXiv:2310.01405, 2023
Pith/arXiv arXiv 2023
-
[34]
Diff-in-means concept editing is worst-case optimal, 2023
Nora Belrose. Diff-in-means concept editing is worst-case optimal, 2023. URL https: //blog.eleuther.ai/diff-in-means/. Accessed: 2026-04-29
2023
-
[35]
Inference- time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference- time intervention: Eliciting truthful answers from a language model. InAdv. Neural Inform. Process. Syst. (NeurIPS), 2023
2023
-
[36]
Linear representations of sentiment in large language models.arXiv preprint arXiv:2310.15154, 2023
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. Linear representations of sentiment in large language models.arXiv preprint arXiv:2310.15154, 2023
Pith/arXiv arXiv 2023
-
[37]
Improving instruction-following in language models through activation steering
Alessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz, and Besmira Nushi. Improving instruction-following in language models through activation steering. InInt. Conf. Learn. Rep. (ICLR), 2025
2025
-
[38]
The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Pith/arXiv arXiv 2024
-
[39]
Qwen2.5 technical report.arxiv preprint arXiv:2412.15115, 2024
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li,...
Pith/arXiv arXiv 2024
-
[40]
Openai usage policies, 2025
OpenAI. Openai usage policies, 2025. URL https://openai.com/policies/ usage-policies. Accessed: 2026-05-19
2025
-
[41]
Exponential moving average of weights in deep learning: Dynamics and benefits.Transactions on Machine Learning Research Journal, pages 1–27, 2024
Daniel Morales Brotons, Thijs V ogels, and Hadrien Hendrikx. Exponential moving average of weights in deep learning: Dynamics and benefits.Transactions on Machine Learning Research Journal, pages 1–27, 2024
2024
-
[42]
Rethinking safety in llm fine-tuning: An optimization perspective
Minseon Kim, Jin Myung Kwak, Lama Alssum, Bernard Ghanem, Philip Torr, David Krueger, Fazl Barez, and Adel Bibi. Rethinking safety in llm fine-tuning: An optimization perspective. InSecond Conference on Language Modeling, 2025
2025
-
[43]
Pku-saferlhf: Towards multi-level safety alignment for llms with human preference
Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Alex Qiu, Jiayi Zhou, Kaile Wang, Boxun Li, et al. Pku-saferlhf: Towards multi-level safety alignment for llms with human preference. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 31983–32016, 2025
2025
-
[44]
Fine-tuning aligned language models compromises safety, even when users do not intend to! In Int
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! In Int. Conf. Learn. Rep. (ICLR), 2024
2024
-
[45]
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. Adv. Neural Inform. Process. Syst. (NeurIPS), 37:55005–55029, 2024
2024
-
[46]
Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. InConf. North Am. Chapter Assoc. Comput. Linguistics (NAACL), pages 5377–5400, 2024
2024
-
[47]
The art of saying no: Contextual noncompliance in language models.Adv
Faeze Brahman, Sachin Kumar, Vidhisha Balachandran, Pradeep Dasigi, Valentina Pyatkin, Abhilasha Ravichander, Sarah Wiegreffe, Nouha Dziri, Khyathi Chandu, Jack Hessel, et al. The art of saying no: Contextual noncompliance in language models.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:49706–49748, 2024. 13
2024
-
[48]
Measuring massive multitask language understanding.Int
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.Int. Conf. Learn. Rep. (ICLR), 2021
2021
-
[49]
Aligning ai with shared human values.Int
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning ai with shared human values.Int. Conf. Learn. Rep. (ICLR), 2021
2021
-
[50]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168, 2021
Pith/arXiv arXiv 2021
-
[51]
Instruction-following evaluation for large language models.arXiv preprint arXiv:2311.07911, 2023
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models.arXiv preprint arXiv:2311.07911, 2023
Pith/arXiv arXiv 2023
-
[52]
Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374, 2021
Pith/arXiv arXiv 2021
-
[53]
Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, et al. Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues. InAnn. Meet. Assoc. Comput. Linguistics (ACL), pages 7421–7454, 2024
2024
-
[54]
A survey on llm-as-a-judge.The Innovation, 2024
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge.The Innovation, 2024
2024
-
[55]
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. Enhancing chat language models by scaling high-quality instructional conversations. InConf. Empir . Methods Nat. Lang. Process. (EMNLP), pages 3029–3051, 2023
2023
-
[56]
The geometry of refusal in large language models: Concept cones and representational independence
Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, and Johannes Gasteiger. The geometry of refusal in large language models: Concept cones and representational independence. InInternational Conference on Machine Learning, pages 66945–66970. PMLR, 2025
2025
-
[57]
The hidden dimensions of llm alignment: A multi-dimensional analysis of orthogonal safety directions
Wenbo Pan, Zhichao Liu, Qiguang Chen, Xiangyang Zhou, Yu Haining, and Xiaohua Jia. The hidden dimensions of llm alignment: A multi-dimensional analysis of orthogonal safety directions. InInt. Conf. Mach. Learn. (ICML), pages 47697–47716. PMLR, 2025
2025
-
[58]
Differentiated directional intervention: A framework for evading llm safety alignment
Peng Zhang and Peijie Sun. Differentiated directional intervention: A framework for evading llm safety alignment. InAAAI, volume 40, pages 38102–38110, 2026
2026
-
[59]
Leheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang, Xiaohao Liu, Zhenkai Liang, Xiang Wang, An Zhang, and Tat-Seng Chua. Alphasteer: Learning refusal steering with principled null-space constraint.arXiv preprint arXiv:2506.07022, 2025
arXiv 2025
-
[60]
Analysing the generalisation and reliability of steering vectors.Adv
Daniel Tan, David Chanin, Aengus Lynch, Brooks Paige, Dimitrios Kanoulas, Adrià Garriga- Alonso, and Robert Kirk. Analysing the generalisation and reliability of steering vectors.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:139179–139212, 2024
2024
-
[61]
Or-bench: An over-refusal benchmark for large language models
Justin Cui, Wei-Lin Chiang, Ion Stoica, and Cho-Jui Hsieh. Or-bench: An over-refusal benchmark for large language models. InInt. Conf. Mach. Learn. (ICML), pages 11515–11542. PMLR, 2025
2025
-
[62]
Universal jailbreak suffixes are strong attention hijackers.arXiv preprint arXiv:2506.12880, 2025
Matan Ben-Tov, Mor Geva, and Mahmood Sharif. Universal jailbreak suffixes are strong attention hijackers.arXiv preprint arXiv:2506.12880, 2025
arXiv 2025
-
[63]
Flipattack: Jailbreak llms via flipping
Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, and Bryan Hooi. Flipattack: Jailbreak llms via flipping. InInt. Conf. Mach. Learn. (ICML), pages 38623–38663. PMLR, 2025. 14
2025
-
[64]
Many-shot jailbreaking.Adv
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, et al. Many-shot jailbreaking.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:129696–129742, 2024
2024
-
[65]
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. InInt. Conf. Mach. Learn. (ICML), 2024
2024
-
[66]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7b.arXiv preprint arXiv:23...
Pith/arXiv arXiv 2023
-
[67]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
Pith/arXiv arXiv 2024
-
[68]
Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. Gemma 2: Improving open language models at a practical size.arXiv preprint arXiv:2408.00118, 2024
Pith/arXiv arXiv 2024
-
[69]
Hashimoto
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca, 2023
2023
-
[70]
Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Adv
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:8093–8131, 2024
2024
-
[71]
What are some good books on Roman history?
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Adv. Neural Inform. Process. Syst. (NeurIPS), 37:47094–47165, 2024. A Internal Representations of Harmfulness and R...
arXiv 2024
This paper was first reviewed by grok-4.5 on July 12, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.