REVIEW 13 cited by
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in generative AI have enabled ubiquitous access to large language models (LLMs). Empowered by their exceptional capabilities to understand and generate human-like text, these models are being increasingly integrated into our society. At the same time, there are also concerns on the potential misuse of this powerful technology, prompting defensive measures from service providers. To overcome such protection, jailbreaking prompts have recently emerged as one of the most effective mechanisms to circumvent security restrictions and elicit harmful content originally designed to be prohibited. Due to the rapid development of LLMs and their ease of access via natural languages, the frontline of jailbreak prompts is largely seen in online forums and among hobbyists. To gain a better understanding of the threat landscape of semantically meaningful jailbreak prompts, we systemized existing prompts and measured their jailbreak effectiveness empirically. Further, we conducted a user study involving 92 participants with diverse backgrounds to unveil the process of manually creating jailbreak prompts. We observed that users often succeeded in jailbreak prompts generation regardless of their expertise in LLMs. Building on the insights from the user study, we also developed a system using AI as the assistant to automate the process of jailbreak prompt generation.
Forward citations
Cited by 13 Pith papers
-
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
EoBench's 19-type, 65,778-item benchmark shows LLMs follow false in-context beliefs more when phrased as imperatives, child-directed speech, formal assertions, or authority appeals, while instruction-tuning and larger...
-
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
A bi-level adversarial training method where a hypernetwork generates malicious LoRA patches to attack the defender, and the defender learns to nullify them, improves tamper resistance across ten open-weight LLMs with...
-
Privacy and Security Threat for OpenAI GPTs
A large-scale study finds that over 98.8% of sampled OpenAI custom GPTs leak their system instructions to crafted adversarial prompts, and hundreds of GPTs transmit user conversation data to third parties.
-
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
ICE decomposes a harmful prompt into hierarchy fragments and semantic hints, wraps them in a fake reasoning task, and reports state-of-the-art single-query jailbreak success, plus a new dual-scenario dataset.
-
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
SAIF is a proposed framework that generates multimodal test prompts from a risk taxonomy, jailbreak tricks, and prompt styles to evaluate generative AI risks in the public sector.
-
Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models
JMLLM, a hybrid obfuscation framework, raises jailbreak success rates across text, image, and speech inputs of multimodal LLMs while using fewer queries than prior methods.
-
The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models
A multilingual jailbreak benchmark on four proprietary LLMs finds language-dependent safety gaps and identifies a two-sided debate prompt as the most effective attack component.
-
`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs
A voice jailbreak that buries a forbidden question between benign prompts reportedly succeeds against Gemini 67 to 93 percent of the time, but the metric comes from the target model judging itself and is not reliable.
-
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
An iterative LLM-based prompt evaluator blocked 100% of the Best-of-N jailbreaking paper's released successful prompts and 99.8% of a fresh replication, with false-positive rates near zero.
-
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
A simple gradient-free 'directional gradient approximation' attack makes LLM time series forecasters degrade more than equivalent random noise, across GPT-3.5, GPT-4, LLaMa, Mistral, TimeGPT, and TimeLLM.
-
A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models
A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.
-
Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages
Relabeling hate speech as metaphor pairs (red/green, summer/winter) in prompts raises Llama2's F1 on a 500-item Bengali subsample to 95.89, though the gain is reported without matched test-set comparisons or error bars.
-
Preventing Jailbreak Prompts as Malicious Tools for Cybercriminals: A Cyber Defense Perspective
A structured survey of jailbreak prompts and layered defenses for large language models, with six illustrative case studies and no empirical evaluation.
Discussion (0). Continue with ORCID to comment.