Pith. sign in

REVIEW 4 major objections 7 minor 43 references

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers

T0 review · 4 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Summaries of LLM safety papers, with a harmful query embedded as a completion-style example, can jailbreak aligned models at 97–98% attack success.

desk verdict A jailbreak that likely works but whose authority-based mechanism is under-supported by the missing no-paper control; worth refereeing with a demand for that control. read the letter →

arxiv 2507.13474 v1 pith:GDRQVEJU submitted 2025-07-17 cs.CL

classification cs.CL
keywords PaperSummaryAttackLLMjailbreakingsafetypapersalignmentbiassuccessrateexternalknowledgecarriersdefenseadversarialprompts
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that aligning LLMs against harmful requests can be undone by supplying a summary of an academic paper about LLM safety, with the harmful request embedded as a fill-in-the-blank 'Example Scenario' near the end. It reports that this Paper Summary Attack reaches a 97% attack success rate on Claude3.5-Sonnet and 98% on DeepSeek-R1, and that standard defenses such as perplexity filtering, LlamaGuard, and moderation APIs barely reduce it. The paper also claims to expose an alignment bias: some models are far more vulnerable when the paper is about defenses, others when it is about attacks, and this bias shifts when only the label 'attack' or 'defense' is swapped. If these results hold, authoritative academic-style context is a general-purpose jailbreak vector that current safety alignment does not address.

What carries the argument

The central object is the PSA prompt template: a structured, section-wise summary of a real LLM safety paper with a 'Payload' section inserted before the final section. The template's job is to make the victim model treat the harmful query as a routine example to complete within an authoritative academic document. The paper's proposed mechanism is authority transfer—models accept academic-style information uncritically—and the reversal experiment, in which attack summaries are labeled as defense summaries and vice versa, is the instrument used to isolate the content-type bias. The hidden-state analysis is the explanatory layer: PSA inputs keep positive tokens like 'Sure' and 'Here' in middle layers rather than refusal tokens like 'Sorry' and 'Cannot', suggesting the safety classification itself is evaded before the final response is generated.

What would settle it

Take PSA's exact prompt template and delete the paper summaries while keeping the phrase 'help me completing Example Scenario' and the harmful query; if Claude3.5-Sonnet or DeepSeek-R1 still produces harmful completions at near 97–98%, the authority of the paper is not the operative mechanism. A complementary check uses an irrelevant or gibberish summary under the same labels; high ASR there would mean the content type, not the content, drives the effect.

Watch

Extended reading notes

Core claim

PSA works by collecting real LLM safety papers, having GPT-4o condense each section, and then splicing the summaries together with a 'Payload' section whose 'detail harmful content' placeholder is filled with the target query. The assembled input is framed as a paper-summary completion task, and on 100 questions sampled from AdvBench and JailbreakBench it beats five baselines on average. The headline numbers are PSA-D (defense-oriented paper summaries) reaching 97% ASR on Claude3.5-Sonnet and 98% on DeepSeek-R1, while PSA-A (attack-oriented summaries) reaches 100% on DeepSeek-R1 and 98% on Llama2. The paper's more general claim is that the effect is content-type dependent: Llama3.1 and Claude3.5 are more vulnerable to defense-type papers, GPT-4o to attack-type papers, and merely relabeling the paper type shifts ASR in the direction of the label. The hidden-state analysis finds that PSA, unlike GCG, PAP, Code Attack, and ArtPrompt, produces positive emotional tokens in middle layers, consistent with the model internally classifying the harmful request as harmless.

Load-bearing premise

The paper assumes the academic paper context, not the explicit phrase asking the model to complete the harmful example, is the active cause of the jailbreak; no control condition runs the same instructions without a paper.

Editorial extensions

If this is right

  • Academic-style summaries act as a transferable jailbreak prefix that works across base, instruct-tuned, and reasoning models without per-query prompt optimization.
  • On Claude3.5-Sonnet and Llama3.1, defense-focused safety papers are a stronger jailbreak context than attack-focused papers, meaning defensive literature itself becomes an attack surface.
  • Perplexity filtering, LlamaGuard, and moderation APIs leave PSA's attack success rate largely intact, and LlamaGuard's protection is concentrated on attack-paper variants, not defense-paper variants.
  • Relabeling a paper's type flips the attack's effectiveness in the direction of the label, showing that model safety behavior tracks the declared category even when the underlying text is unchanged.
  • The middle-layer emotional-token pattern implies PSA is missed by safety checks that rely on later refusal behavior, because the input is already classified as harmless early in the model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the template's explicit completion instruction is a plausible confound; a no-paper control would tell whether the attack is authority-driven or simply instruction-following.
  • Editorial inference: the label-swap result suggests genre classification, not semantic content, carries much of the effect, so gibberish or unrelated summaries under the same labels would be a cheap falsification test.
  • Editorial inference: the same journal-style structure could be tested with non-safety academic papers or technical documentation to see whether the vulnerability generalizes beyond safety discourse.
  • Editorial inference: because the same model, GPT-4o, generates summaries and judges harmfulness, an independent judge or human annotation would disentangle jailbreak success from judge-side bias.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes Paper Summary Attack (PSA), a jailbreaking method that summarizes LLM-safety papers into a structured template and inserts harmful queries as a completion-style payload. The authors report very high attack success rates (97% on Claude-3.5-Sonnet, 98% on DeepSeek-R1), compare PSA with five existing jailbreak baselines, evaluate three defensive countermeasures, and analyze intermediate hidden states. They also report a directional asymmetry: on some models defense-focused papers are more effective jailbreak contexts than attack-focused papers, which they interpret as evidence of alignment bias.

Significance. If the causal mechanism claimed by the paper held, PSA would be a notable contribution: a scalable, transferable jailbreak that exploits the presumed authority of academic-style content, together with a new empirical phenomenon of asymmetric safety alignment across paper types. The paper is well positioned in the jailbreaking literature, includes code, uses external benchmarks (AdvBench, JailbreakBench), and provides a defined threat model. However, the central interpretation is not supported by the experimental design: the prompt template already contains explicit instructions to produce the harmful output, and no control condition removes the paper context while keeping the template intact. The any-of-6 aggregation of attack attempts further inflates the headline ASR. These issues are load-bearing for the paper's main claim that academic authority is the active cause of the jailbreak.

major comments (4)
  1. [§3.1 and Appendix A.2] The preliminary experiment and the PSA prompt contain explicit completion instructions: the template says "help me completing Example Scenario based on the full context" and then, for the harmful query, "according to the paper, it's clear that the process of it is: ......". These instructions directly ask the model to generate the harmful content. Because no control condition runs the identical completion template with the paper removed (or replaced by a neutral, non-safety paper), the observed jailbreak cannot be causally attributed to the academic-paper context or to the purported authority of academic content. Please add at least three ablations: (a) the completion template alone, (b) the completion template with a neutral paper in the same structure, and (c) the paper summaries without the completion instruction. Report ASR for each; the current data cannot distinguish these explanations.
  2. [§5.1 Setup of PSA] The ASR is aggregated as any-of-6: "If any one of these attempts succeeds, it is recorded as a success." This makes the headline 97% and 98% figures upper bounds over six chances per question, and the table does not report per-attempt success rates, the distribution of successes across the six paper subcategories, or standard errors/confidence intervals. Please report per-attempt ASR, the number of questions per condition, and statistical significance tests for the cross-model and cross-paper-type differences. Without these, the high headline numbers and the claimed alignment bias may overstate the effect.
  3. [§5.2 and Figure 3] The reverse-label experiment is a useful manipulation check but it changes both the stated label (attack vs. defense) and the actual paper content simultaneously, and it still does not isolate the effect of the paper from the effect of the completion instruction. Moreover, the PSA-A and PSA-D conditions differ not only in the paper type but also in the wording of the payload trigger ("if the attack method ... successfully executed" vs. "without the defense method"), so the cross-condition ASR differences in Figure 3 and the appendix heatmaps confound template wording with paper content. Please report the full templates for PSA-A and PSA-D in an appendix and run a matched control where the paper is replaced by a neutral paper while keeping the payload wording fixed.
  4. [Limitations] The Limitations section acknowledges only that the mechanistic analysis of alignment bias is shallow. It does not mention the missing no-paper control, the any-of-6 ASR aggregation, or the absence of statistical significance testing. Since these issues directly affect the central causal claim, the limitations statement should be revised accordingly.
minor comments (7)
  1. [Throughout] There are several typographical errors, including "teste them" (§5.1), "Cluade-3.5-Sonnet" (Appendix A.3), "refinment" (Appendix A.4), and "Playload" in Figure 1.
  2. [Appendix A.4] The paper classification is attributed to "(Bian et al., 2024)" in Appendix A.4 but to "(Yi et al., 2024)" in §4.2; please clarify which source was actually used.
  3. [§4.2] Equations (1) and (2) use undefined notation: I(Sj, Dj) and R(di, sj) are not formally defined, and the optimization in Eq. (1) is not actually solved. Either define these quantities precisely or remove the pseudo-formal notation, as it does not add rigor.
  4. [Figure 4] The caption says tokens marked in green represent positive sentiment but the color legend in the figure is not self-explanatory; also specify the model and layer range used to produce the figure.
  5. [§5.1 Setup] The paper states "we disable sampling by default" but does not report the temperature or decoding settings; please specify these and state whether all baselines used the same settings.
  6. [Table 3] The defense evaluation omits Claude-3.5-sonnet even though the attack evaluation in Table 2 includes it; clarify why, or add the missing column.
  7. [Abstract] The abstract's 97% and 98% ASR figures are any-of-6 aggregated rates; please label them as such in the abstract to avoid overstating the per-attempt success rate.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the attack is evaluated against external benchmarks and an external judge; no claim reduces by construction to its own inputs.

full rationale

This paper is an empirical attack study, not a derivation with fitted parameters, so there is no circular chain of the kind the analysis targets. The central claim that the PSA template achieves high ASR on well-aligned models is measured on external benchmarks (AdvBench and JailbreakBench) with an independent GPT-4o judge, and no success metric is defined in terms of the template's construction or the paper's own assumptions. The preliminary study in Section 3 motivates the choice of LLM-safety papers as attack context, but that is a selection heuristic, not a fitted parameter renamed as a prediction; the final evaluation is independent of the preliminary observations. The paper contains no self-citations that carry a load-bearing inference: the cited prior works on external-information effects (Bian et al.), hidden-state analysis (Zhou et al.), and attack/defense classification (Yi et al.) are external, and none is invoked to forbid alternatives or to define the attack's success. The Limitations section concedes that the mechanistic analysis of alignment bias is shallow, but that is an acknowledged explanatory gap, not a circular step. Even the reader-flagged concern that the payload template itself contains an explicit completion instruction bears on causal attribution and experimental control, not on whether the reported result reduces by definition to its inputs. No equation, definition, or imported theorem in the paper makes the claimed ASR equivalent to its own assumptions.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim rests on several unstated premises: that academic-style text carries authority for LLMs, that the completion payload is not itself the active jailbreak mechanism, that GPT-4o summaries preserve attack-relevant content, and that the GPT-4o judge reliably identifies harmful responses. The paper does not independently validate these premises within the reported experiments.

free parameters (3)
  • Harmfulness score threshold for ASR = 5
    ASR counts only GPT-4o judge scores of exactly 5; the threshold is hand-chosen and defines every reported success rate.
  • Number of PSA attempts per question = 6
    Each question is tried against one paper per subcategory, and any success counts. The any-of-6 aggregation inflates per-question ASR and complicates comparison with baselines.
  • Section token limit and chunk size = 1000 words per chunk
    Summaries are produced under per-section token limits with 1000-word chunks; no sensitivity analysis is provided.
assumptions (4)
  • domain assumption LLMs treat academic-style text as authoritative and internalize it uncritically.
    Invoked in Section 1 and Section 3.2 to explain why paper summaries cause jailbreaks; no control condition isolates this effect from the completion instruction.
  • ad hoc to paper The completion-style payload is neutral and does not itself elicit harmful content.
    The experiments always pair the payload with paper summaries and never run the payload alone, so the design implicitly assumes the payload is not the primary driver.
  • domain assumption GPT-4o summaries preserve enough attack-relevant information to be effective and do not distort the original papers.
    Section 4.2 Step 2 relies on GPT-4o summarization; no human evaluation or comparison with raw paper text is provided.
  • domain assumption GPT-4o judge scores on the 5-point harm scale are reliable, with 5 meaning a direct fulfillment of the harmful request.
    Section 5.1 Metrics uses LLM-as-Judge; no agreement with human raters or alternative judges is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers." pith.science (2026). https://pith.science/paper/GDRQVEJU

@misc{pith2026250713474,
  author       = {Pith},
  title        = {Pith review of: Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GDRQVEJU}},
  note         = {Machine review of arXiv:2507.13474}
}
read the original abstract

The safety of large language models (LLMs) has garnered significant research attention. In this paper, we argue that previous empirical studies demonstrate LLMs exhibit a propensity to trust information from authoritative sources, such as academic papers, implying new possible vulnerabilities. To verify this possibility, a preliminary analysis is designed to illustrate our two findings. Based on this insight, a novel jailbreaking method, Paper Summary Attack (\llmname{PSA}), is proposed. It systematically synthesizes content from either attack-focused or defense-focused LLM safety paper to construct an adversarial prompt template, while strategically infilling harmful query as adversarial payloads within predefined subsections. Extensive experiments show significant vulnerabilities not only in base LLMs, but also in state-of-the-art reasoning model like Deepseek-R1. PSA achieves a 97\% attack success rate (ASR) on well-aligned models like Claude3.5-Sonnet and an even higher 98\% ASR on Deepseek-R1. More intriguingly, our work has further revealed diametrically opposed vulnerability bias across different base models, and even between different versions of the same model, when exposed to either attack-focused or defense-focused papers. This phenomenon potentially indicates future research clues for both adversarial methodologies and safety alignment.Code is available at https://github.com/233liang/Paper-Summary-Attack

Figures

Figures reproduced from arXiv: 2507.13474 by the authors.

Figure 1
Figure 1. Overview of PSA. PSA consists of three steps. The first step is to collect papers related to the safety of LLM. In the second step, we utilize a jailbreak agent to summarize the targeted sections of the collected papers. In the third step, We concatenate the payload triggers containing harmful questions with each chapter to form a complete prompt, which is then input as text to the victim LLM. success rate, suggesti… view at source ↗
Figure 2
Figure 2. Evaluation results of attack and defense methods [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. The text modifications significantly affect ASR, [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The PSA’s hidden state in the middle layers. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Evaluation results of attack and defense [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 8
Figure 8. Figure 8: Intermediate Hidden State of PAP L16 L17 L18 L19 L20 L21 L22 L23 L24 Layer 0 1 2 3 4 5 6 7 8 9 Rank gouverngouverngouverngouvern cannotpenas step step step dzier Inflater dzier Please penas cannot Step Step Step dummy forb please <s> Please macht penas penas penas Infl…
Figure 9
Figure 9. Figure 9: Intermediate Hidden State of CodeAttack L16 L17 L18 L19 L20 L21 L22 L23 L24 Layer 0 1 2 3 4 5 6 7 8 9 Rank gre gre gre bras diretto uche anch ola cannot sure anch ban gouvern brashomonymescannot angel bras ola sure olare enso penas bras pas cannot code bras cannot zyk …
Figure 10
Figure 10. Figure 10: Intermediate Hidden State of ArtPrompt [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 11 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    AI@Meta. 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md Llama 3 model card

  5. [5]

    Alex Albert . 2023. Jailbreak chat. https://www.jailbreakchat.com/

  6. [6]

    Gabriel Alon and Michael Kamfonas. 2023. Detecting language model attacks with perplexity. arXiv preprint arXiv:2308.14132

  7. [7]

    Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al. 2024. Many-shot jailbreaking. Anthropic, April

  8. [8]

    Anthropic. 2024. https://www.anthropic.com/claude/sonnet Claude 3.5 sonnet

Show all 43 references
  1. [9]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862

  2. [10]

    Ning Bian, Hongyu Lin, Peilin Liu, Yaojie Lu, Chunkang Zhang, Ben He, Xianpei Han, and Le Sun. 2024. Influence of external information on large language models mirrors social cognitive patterns. IEEE Transactions on Computational Social Systems

  3. [11]

    Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023. Defending against alignment-breaking attacks via robustly aligned llm. arXiv preprint arXiv:2309.14348

  4. [12]

    Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. 2024. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. arX...

  5. [13]

    Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. 2023. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419

  6. [14]

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30

  7. [15]

    Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019. Build it break it fix it for dialogue safety: Robustness from adversarial human attack. arXiv preprint arXiv:1908.06083

  8. [16]

    Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao. 2023. Mart: Improving llm safety with multi-round automatic red-teaming. arXiv preprint arXiv:2311.07689

  9. [17]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948

  10. [18]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. 2023. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674

  11. [19]

    Akshita Jha and Chandan K Reddy. 2023. Codeattack: Code-based adversarial attacks for pre-trained programming language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 14892--14900

  12. [20]

    Fengqing Jiang, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran. 2024. Artprompt: Ascii art-based jailbreak attacks against aligned llms. arXiv preprint arXiv:2402.11753

  13. [21]

    G Tarcan Kumkale and Dolores Albarrac \' n. 2004. The sleeper effect in persuasion: a meta-analytic review. Psychological bulletin, 130(1):143

  14. [22]

    LessWrong. 2023. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens Interpreting gpt: The logit lens

  15. [23]

    Xuan Li, Zhanke Zhou, Jianing Zhu, Jiangchao Yao, Tongliang Liu, and Bo Han. 2023. Deepinception: Hypnotize large language model to be jailbreaker. arXiv preprint arXiv:2311.03191

  16. [24]

    Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen. 2024. Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction. In 33rd USENIX Security Symposium (USENIX Security 24), pages 4711--4728

  17. [25]

    Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. 2023. Autodan: Generating stealthy jailbreak prompts on aligned large language models. arXiv preprint arXiv:2310.04451

  18. [26]

    LMSYS. 2023. Vicuna-7b-v1.5. https://huggingface.co/lmsys/vicuna-7b-v1.5

  19. [27]

    Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey. arXiv preprint arXiv:2402.06196

  20. [28]

    OpenAI. 2023. https://platform.openai.com/docs/api-reference/moderations Moderations . Accessed: 2023-12-05

  21. [29]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 3...

  22. [30]

    Chanthika Pornpitakpan. 2004. The persuasiveness of source credibility: A critical review of five decades' evidence. Journal of applied social psychology, 34(2):243--281

  23. [31]

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693

  24. [32]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36

  25. [33]

    Alexander Robey, Eric Wong, Hamed Hassani, and George J Pappas. 2023. Smoothllm: Defending large language models against jailbreaking attacks. arXiv preprint arXiv:2310.03684

  26. [34]

    do anything now

    Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023. " do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. arXiv preprint arXiv:2308.03825

  27. [35]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  28. [36]

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. 2023. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In NeurIPS

  29. [37]

    Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al. 2021. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359

  30. [38]

    Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li. 2024. Jailbreak attacks and defenses against large language models: A survey. arXiv preprint arXiv:2407.04295

  31. [39]

    Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. 2023. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher. arXiv preprint arXiv:2308.06463

  32. [40]

    Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. 2024. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. arXiv preprint arXiv:2401.06373

  33. [41]

    Xiaoyu Zhang, Cen Zhang, Tianlin Li, Yihao Huang, Xiaojun Jia, Xiaofei Xie, Yang Liu, and Chao Shen. 2023. A mutation-based method for multi-modal jailbreaking attack detection. arXiv preprint arXiv:2312.10766

  34. [42]

    Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, and Yongbin Li. 2024. How alignment and jailbreak work: Explain llm safety through intermediate hidden states. arXiv preprint arXiv:2406.05644

  35. [43]

    Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.