Pith. sign in

REVIEW 11 cited by

Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.09127 v2 pith:CKLFA5CB submitted 2023-11-15 cs.CR cs.AIcs.LG

classification cs.CRcs.AIcs.LG
keywords systempromptsjailbreakjailbreakingattackgpt-4vpotentialsuccess
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing work on jailbreak Multimodal Large Language Models (MLLMs) has focused primarily on adversarial examples in model inputs, with less attention to vulnerabilities, especially in model API. To fill the research gap, we carry out the following work: 1) We discover a system prompt leakage vulnerability in GPT-4V. Through carefully designed dialogue, we successfully extract the internal system prompts of GPT-4V. This finding indicates potential exploitable security risks in MLLMs; 2) Based on the acquired system prompts, we propose a novel MLLM jailbreaking attack method termed SASP (Self-Adversarial Attack via System Prompt). By employing GPT-4 as a red teaming tool against itself, we aim to search for potential jailbreak prompts leveraging stolen system prompts. Furthermore, in pursuit of better performance, we also add human modification based on GPT-4's analysis, which further improves the attack success rate to 98.7\%; 3) We evaluated the effect of modifying system prompts to defend against jailbreaking attacks. Results show that appropriately designed system prompts can significantly reduce jailbreak success rates. Overall, our work provides new insights into enhancing MLLM security, demonstrating the important role of system prompts in jailbreaking. This finding could be leveraged to greatly facilitate jailbreak success rates while also holding the potential for defending against jailbreaks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm

    cs.CR 2025-09 conditional novelty 7.0 of 10

    A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.

  2. SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

    cs.CR 2025-10 conditional novelty 6.0 of 10

    A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.

  3. Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A small set of attention heads in LVLMs flags malicious prompts during the first token; a logistic-regression detector built on them reduces jailbreak success to 1-5%.

  4. VLSBench: Unveiling Visual Leakage in Multimodal Safety

    cs.CR 2024-11 conditional novelty 6.0 of 10

    The paper shows existing multimodal safety benchmarks leak harmful image content into text queries (VSIL), and introduces VLSBench, a 2.2k-pair leakless benchmark on which textual alignment fails and multimodal alignm...

  5. Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models

    cs.CL 2024-11 conditional novelty 6.0 of 10

    An automated agent can jailbreak GPT-4o and other vision-language models using only individually safe images and benign-sounding prompts, escalating responses to harmful content.

  6. The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense

    cs.CR 2024-11 conditional novelty 6.0 of 10

    Near-perfect jailbreak defenses for vision-language models are mostly over-refusal, and the two standard ways of scoring jailbreaks agree only at chance level.

  7. One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs

    cs.CR 2025-05 conditional novelty 5.0 of 10

    ArrAttack fine-tunes a judge on the SmoothLLM defense, uses it to filter rewriting-attack data, and trains a generator that produces jailbreak prompts transferring across defenses.

  8. SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector

    cs.AI 2025-01 conditional novelty 5.0 of 10

    SAIF is a proposed framework that generates multimodal test prompts from a risk taxonomy, jailbreak tricks, and prompt styles to evaluate generative AI risks in the public sector.

  9. Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense

    cs.CR 2025-01 reject novelty 5.0 of 10

    Layer-AdvPatcher edits 'toxic' transformer layers using self-generated harmful examples to block jailbreaks, but its reported attack-success rates worsen on several benchmarks.

  10. Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A survey that taxonomizes multimodal jailbreak attacks and defenses into four lifecycle levels (input, encoder, generator, output) across Any-to-Text, Any-to-Vision, and Any-to-Any generative models.

  11. When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs

    cs.CV 2025-02 conditional novelty 3.0 of 10

    A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.

Pith tools