REVIEW 42 cited by
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Most traditional AI safety research has approached AI models as machines and centered on algorithm-focused attacks developed by security experts. As large language models (LLMs) become increasingly common and competent, non-expert users can also impose risks during daily interactions. This paper introduces a new perspective to jailbreak LLMs as human-like communicators, to explore this overlooked intersection between everyday language interaction and AI safety. Specifically, we study how to persuade LLMs to jailbreak them. First, we propose a persuasion taxonomy derived from decades of social science research. Then, we apply the taxonomy to automatically generate interpretable persuasive adversarial prompts (PAP) to jailbreak LLMs. Results show that persuasion significantly increases the jailbreak performance across all risk categories: PAP consistently achieves an attack success rate of over $92\%$ on Llama 2-7b Chat, GPT-3.5, and GPT-4 in $10$ trials, surpassing recent algorithm-focused attacks. On the defense side, we explore various mechanisms against PAP and, found a significant gap in existing defenses, and advocate for more fundamental mitigation for highly interactive LLMs
Forward citations
Cited by 42 Pith papers
-
LLMs Encode Harmfulness and Refusal Separately
LLMs encode a separate internal harmfulness direction, distinct from the refusal direction, which is more robust to jailbreaks and adversarial finetuning.
-
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
Evaluating LLM safety with one canonical prompt understates unsafe behavior; across five meaning-preserving reformulations, 5-13% of safe-on-canonical seeds become unsafe, and the union exceeds the worst single form f...
-
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
RoguePrompt, a Vigenère+ROT13 self-reconstruction jailbreak, achieves 70.18% execution@3 and 93.93% bypass@3 across GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro on 313 StrongREJECT prompts.
-
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
LLMs sometimes regain safe behavior after multi-turn jailbreaks, and this recovery can be measured with turn-level safety trajectories and metrics such as misalignment length and recovery duration.
-
Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion
Vision-language models are more easily persuaded by multimodal messages than by text alone, especially in adversarial misinformation settings.
-
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
Special tokens that structure LLM conversations can be injected and swapped for lookalike words to bypass both built-in safety and external content filters.
-
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models
Seed2Harvest expands 1,000 human adversarial prompts into 27,650 LLM-generated variants that keep roughly comparable unsafe-image trigger rates and add hundreds of new geographic contexts.
-
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers
Summaries of LLM safety papers, paired with a completion-style payload containing a harmful query, jailbreak aligned LLMs at high reported success rates and expose a defense paper versus attack paper bias.
-
MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
MGC, a two-stage compiler framework, generates functional malware by decomposing malicious intents into benign-appearing MDIR components that strong aligned LLMs will implement, bypassing safety alignment.
-
HauntAttack: When Attack Follows Reasoning as a Shadow
HauntAttack embeds harmful instructions into reasoning-question conditions and reports a 70% average attack success rate across 11 large reasoning models, outperforming prior jailbreak baselines.
-
Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
A component-based, genetically optimized jailbreak framework reports over 90% success on Claude-3.5 and strong cross-model transferability.
-
Position: Adversarial ML for LLMs Is Not Making Any Progress
The authors argue that LLM-era adversarial machine learning is less well-defined, harder to solve, and harder to evaluate, so meaningful progress may not be achievable or trackable in the current paradigm.
-
Adversarial Reasoning at Jailbreaking Time
A loss-guided 'reason, verify, search' loop with three LLM modules surpasses prior semantic jailbreak methods on several defended models.
-
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
Scientific-sounding persuasion, using real or fabricated research summaries, reliably increases stereotypical bias and toxicity in multiple commercial LLMs.
-
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
A fairness and robustness audit of four closed-source moderation APIs finds measurable group disparities and shows that LLM-based paraphrasing can bypass unsafe-content flags.
-
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
Auto-RT uses early-terminated exploration plus reward shaping from progressively weakened copies of the target model to automatically discover jailbreak strategies, reporting up to 16.63% higher attack success than baselines.
-
DiffusionAttacker: Diffusion-Driven Prompt Manipulation for LLM Jailbreak
A diffusion-based prompt rewriter that pushes rewritten prompts toward harmless regions of a target model's hidden states achieves higher jailbreak success than existing suffix and template attacks.
-
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
A fine-tuned guideline generator that prepends risk summaries to prompts cuts jailbreak attack success by about 34 percentage points on average across three chatbots without altering the target models.
-
The Mirage of Artificial Intelligence Terms of Use Restrictions
AI model terms of use are likely largely unenforceable because model weights and outputs lack copyright protection, and other legal doctrines do not fill the gap.
-
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
A new attack pipeline, ReG-QA, generates natural, semantically related questions from a toxic seed and jailbreaks aligned LLMs at rates up to 93% on GPT-3.5 and 82% on GPT-4.
-
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
No tested defense, including the paper's own transcript classifier, can fully stop an LLM from giving competent bomb-making instructions under a grey-box attacker.
-
Universal and Context-Independent Triggers for Precise Control of LLM Outputs
A single trained token pair inserted around any target text forces Qwen-2 7B and Llama-3.1 8B to output that text on 54 to 75 percent of unseen prompts.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
InfoFlood: Jailbreaking Large Language Models with Information Overload
InfoFlood claims near-perfect jailbreak success on four frontier LLMs by rewriting harmful queries into verbose academic prose with fake citations, past-tense framing, and ethical disclaimers, without adversarial suffixes.
-
Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models
Emotional prompts make LLMs produce more cognitively complex language than rational prompts, and models systematically differ in their use of Cialdini influence principles across prompt types.
-
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
A distilled open-source attacker, KDA, imitates three jailbreak methods to write diverse attack prompts, and reports higher success and efficiency than each teacher.
-
Self-Instruct Few-Shot Jailbreaking: Decompose the Attack into Pattern and Behavior Learning
A few-shot jailbreak method that combines repeated special-token patterns with self-generated harmful demos to push sample-level attack success near 90% on several open-source LLMs.
-
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
Layer-AdvPatcher edits 'toxic' transformer layers using self-generated harmful examples to block jailbreaks, but its reported attack-success rates worsen on several benchmarks.
-
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models
LLM-Virus uses an evolutionary algorithm with an LLM as crossover, mutation, and fitness operator to evolve jailbreak templates, reporting state-of-the-art attack success on HarmBench and AdvBench.
-
Jailbreaking? One Step Is Enough!
REDA jailbreaks LLMs in one step by framing harmful requests as defensive tasks and using in-context examples, reporting high and transferable attack success rates.
-
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
A new white-box attack edits MLP matrices of safety-aligned open-source LLMs to remove safety-critical transformations, achieving 84.86% average jailbreak success without prompt modification.
-
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
A black-box jailbreak method distributes harmful semantics across text and image inputs and uses heuristic search to induce multimodal LLMs to answer harmful queries.
-
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
An inference-time alignment method using a safety reward model and controlled decoding that reduces jailbreak success rates in multimodal LLMs while preserving utility.
-
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
LLMs are measurably persuadable in morally ambiguous scenarios, with susceptibility varying strongly by model and only slightly with conversation length beyond a few turns.
-
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
Fake authoritative citations matched to the type of harmful request can bypass safety alignment in several commercial and open LLMs.
-
A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination
A new jailbreak framework, HACA, combines atomic text and image attack strategies selected by a cross-modal planner and generates attacks with LLMs and text-to-image models, reaching 95.48% average attack success acro...
-
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
CoopGuard's defer-tempt-analyze-coordinate agents cut reported jailbreak success and raise attacker token costs on the new EMRA benchmark, but the deceptive-rate metric is partly defined by the paper's own scoring rubric.
-
Adversarial Preference Learning for Robust LLM Alignment
APL iteratively trains an attacker to generate adversarial prompt rewrites and a defender to resist them, using the defender's own preference probabilities as the attack signal.
-
Safety Reasoning with Guidelines
Training LLMs to reason through explicit safety guidelines reduces out-of-distribution jailbreak success rates compared to standard refusal training.
-
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
An iterative LLM-based prompt evaluator blocked 100% of the Best-of-N jailbreaking paper's released successful prompts and 99.8% of a fresh replication, with false-positive rates near zero.
-
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
Medical LLMs are highly vulnerable to black-box jailbreaking, and adversarial continual fine-tuning greatly reduces measured jailbreak effectiveness.
-
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
LIAR shows that best-of-N sampling of suffixes from a GPT-2 model jailbreaks several aligned LLMs with low-perplexity prompts and far faster time-to-attack than training-based attacks.
Discussion (0). Continue with ORCID to comment.