REVIEW 8 cited by
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In recent years, large language models (LLMs) have demonstrated notable success across various tasks, but the trustworthiness of LLMs is still an open problem. One specific threat is the potential to generate toxic or harmful responses. Attackers can craft adversarial prompts that induce harmful responses from LLMs. In this work, we pioneer a theoretical foundation in LLMs security by identifying bias vulnerabilities within the safety fine-tuning and design a black-box jailbreak method named DRA (Disguise and Reconstruction Attack), which conceals harmful instructions through disguise and prompts the model to reconstruct the original harmful instruction within its completion. We evaluate DRA across various open-source and closed-source models, showcasing state-of-the-art jailbreak success rates and attack efficiency. Notably, DRA boasts a 91.1% attack success rate on OpenAI GPT-4 chatbot.
Forward citations
Cited by 8 Pith papers
-
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
MOCHA is a benchmark of 10.5K malicious coding prompts, including multi-turn decomposition attacks, showing code LLMs reject these incremental attacks at much lower rates and that fine-tuning on the benchmark improves...
-
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
A bi-level adversarial training method where a hypernetwork generates malicious LoRA patches to attack the defender, and the defender learns to nullify them, improves tamper resistance across ten open-weight LLMs with...
-
Behind the Mask: Benchmarking Camouflaged Jailbreaks in Large Language Models
A benchmark of 500 camouflaged jailbreak prompts finds open-weight LLMs comply with 94% of harmful requests, but the result is confounded by task complexity and an overly permissive compliance metric.
-
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation
Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.
-
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
A distilled open-source attacker, KDA, imitates three jailbreak methods to write diverse attack prompts, and reports higher success and efficiency than each teacher.
-
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
A new white-box attack edits MLP matrices of safety-aligned open-source LLMs to remove safety-critical transformations, achieving 84.86% average jailbreak success without prompt modification.
-
Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents
A survey proposing a source-and-impact taxonomy (input, model, combined; security, privacy, ethics) for threats to LLM-based agents, with feature analysis and four case studies.
-
SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation
A systematization-of-knowledge survey that categorizes LLM privacy risks into training data, prompts, outputs, and agents, and reviews limitations of current mitigations.
Discussion (0). Continue with ORCID to comment.