REVIEW 3 cited by
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as models become increasingly powerful. In this work, we introduce a dataset of jailbreaks where each example can be input in both a single or a multi-turn format. We show that while equivalent in content, they are not equivalent in jailbreak success: defending against one structure does not guarantee defense against the other. Similarly, LLM-based filter guardrails also perform differently depending on not just the input content but the input structure. Thus, vulnerabilities of frontier models should be studied in both single and multi-turn settings; this dataset provides a tool to do so.
Forward citations
Cited by 3 Pith papers
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.
-
Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
Thought-Aligner corrects unsafe intermediate reasoning in LLM agents before actions, raising measured behavioral safety across three benchmarks.
-
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
User-controlled prefill of an LLM's response can force many models to continue with harmful content, but the paper's success-rate metrics may count the attacker's own prefill text.
Discussion (0). Continue with ORCID to comment.