Pith. sign in

REVIEW 5 cited by

Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17336 v2 pith:S7O6ZDVW submitted 2024-03-26 cs.CR cs.CL

classification cs.CRcs.CL
keywords jailbreakpromptsllmsmodelsaccessgenerationlanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in generative AI have enabled ubiquitous access to large language models (LLMs). Empowered by their exceptional capabilities to understand and generate human-like text, these models are being increasingly integrated into our society. At the same time, there are also concerns on the potential misuse of this powerful technology, prompting defensive measures from service providers. To overcome such protection, jailbreaking prompts have recently emerged as one of the most effective mechanisms to circumvent security restrictions and elicit harmful content originally designed to be prohibited. Due to the rapid development of LLMs and their ease of access via natural languages, the frontline of jailbreak prompts is largely seen in online forums and among hobbyists. To gain a better understanding of the threat landscape of semantically meaningful jailbreak prompts, we systemized existing prompts and measured their jailbreak effectiveness empirically. Further, we conducted a user study involving 92 participants with diverse backgrounds to unveil the process of manually creating jailbreak prompts. We observed that users often succeeded in jailbreak prompts generation regardless of their expertise in LLMs. Building on the insights from the user study, we also developed a system using AI as the assistant to automate the process of jailbreak prompt generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 9 citations worldwide. Full citation record

  1. It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief

    cs.CL 2026-07 conditional novelty 6.0 of 10

    EoBench's 19-type, 65,778-item benchmark shows LLMs follow false in-context beliefs more when phrased as imperatives, child-directed speech, formal assertions, or authority appeals, while instruction-tuning and larger...

  2. Privacy and Security Threat for OpenAI GPTs

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A large-scale study finds that over 98.8% of sampled OpenAI custom GPTs leak their system instructions to crafted adversarial prompts, and hundreds of GPTs transmit user conversation data to third parties.

  3. Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion

    cs.CR 2025-05 conditional novelty 5.0 of 10

    ICE decomposes a harmful prompt into hierarchy fragments and semantic hints, wraps them in a fake reasoning task, and reports state-of-the-art single-query jailbreak success, plus a new dual-scenario dataset.

  4. A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models

    cs.IR 2025-07 reject novelty 3.0 of 10

    A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.

  5. Leveraging the Potential of Prompt Engineering for Hate Speech Detection in Low-Resource Languages

    cs.CL 2025-06 conditional novelty 3.0 of 10

    Relabeling hate speech as metaphor pairs (red/green, summer/winter) in prompts raises Llama2's F1 on a 500-item Bengali subsample to 95.89, though the gain is reported without matched test-set comparisons or error bars.

Pith tools