Pith. sign in

REVIEW 9 cited by

Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09091 v2 pith:KYJ7I6UP submitted 2024-02-14 cs.CR cs.AIcs.HC

classification cs.CRcs.AIcs.HC
keywords llmsjailbreakmaliciousattackcluespuzzlerqueryattacks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With the development of LLMs, the security threats of LLMs are getting more and more attention. Numerous jailbreak attacks have been proposed to assess the security defense of LLMs. Current jailbreak attacks primarily utilize scenario camouflage techniques. However their explicitly mention of malicious intent will be easily recognized and defended by LLMs. In this paper, we propose an indirect jailbreak attack approach, Puzzler, which can bypass the LLM's defense strategy and obtain malicious response by implicitly providing LLMs with some clues about the original malicious query. In addition, inspired by the wisdom of "When unable to attack, defend" from Sun Tzu's Art of War, we adopt a defensive stance to gather clues about the original malicious query through LLMs. Extensive experimental results show that Puzzler achieves a query success rate of 96.6% on closed-source LLMs, which is 57.9%-82.7% higher than baselines. Furthermore, when tested against the state-of-the-art jailbreak detection approaches, Puzzler proves to be more effective at evading detection compared to baselines.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VLMInferSlow finds small image perturbations that make black-box VLM APIs produce up to 128% longer responses, sharply increasing latency and energy use.

  2. JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation

    cs.CR 2025-02 conditional novelty 6.0 of 10

    JBShield detects jailbreaks by checking whether a prompt activates both a toxic concept and a jailbreak concept inside an LLM, then steers those concepts to produce a safe refusal.

  3. Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

    cs.CL 2025-08 reject novelty 5.0 of 10

    Sparse autoencoder activation perturbation (SFPF) applied on top of existing jailbreak prompts raises attack success rate on Qwen3-32B, but with no defense evaluation and weak reproducibility.

  4. KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

    cs.CR 2025-02 conditional novelty 5.0 of 10

    A distilled open-source attacker, KDA, imitates three jailbreak methods to write diverse attack prompts, and reports higher success and efficiency than each teacher.

  5. Advancing Jailbreak Strategies: A Hybrid Approach to Exploiting LLM Vulnerabilities and Bypassing Modern Defenses

    cs.CL 2025-06 reject novelty 4.0 of 10

    GCG+PAIR and GCG+WordGame hybrids show mixed attack-success improvements, but methodological flaws, including a questionable GCG loss term and a pre-generated baseline, undermine the paper's main claims.

  6. SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models

    cs.CR 2025-06 conditional novelty 4.0 of 10

    Across 510 HarmBench behaviors and seven attack methods, GPT-4 models show more consistent jailbreak resilience than DeepSeek models, whose vulnerability grows with scale.

  7. `Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMs

    cs.CR 2025-02 reject novelty 4.0 of 10

    A voice jailbreak that buries a forbidden question between benign prompts reportedly succeeds against Gemini 67 to 93 percent of the time, but the metric comes from the target model judging itself and is not reliable.

  8. Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.

  9. A Survey of Attacks on Large Language Models

    cs.CR 2025-05 conditional novelty 1.0 of 10

    A narrative survey that taxonomizes adversarial attacks on LLMs and LLM-based agents into training, inference, and availability/integrity phases with associated defenses.

Pith tools