REVIEW 11 cited by
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Existing work on jailbreak Multimodal Large Language Models (MLLMs) has focused primarily on adversarial examples in model inputs, with less attention to vulnerabilities, especially in model API. To fill the research gap, we carry out the following work: 1) We discover a system prompt leakage vulnerability in GPT-4V. Through carefully designed dialogue, we successfully extract the internal system prompts of GPT-4V. This finding indicates potential exploitable security risks in MLLMs; 2) Based on the acquired system prompts, we propose a novel MLLM jailbreaking attack method termed SASP (Self-Adversarial Attack via System Prompt). By employing GPT-4 as a red teaming tool against itself, we aim to search for potential jailbreak prompts leveraging stolen system prompts. Furthermore, in pursuit of better performance, we also add human modification based on GPT-4's analysis, which further improves the attack success rate to 98.7\%; 3) We evaluated the effect of modifying system prompts to defend against jailbreaking attacks. Results show that appropriately designed system prompts can significantly reduce jailbreak success rates. Overall, our work provides new insights into enhancing MLLM security, demonstrating the important role of system prompts in jailbreaking. This finding could be leveraged to greatly facilitate jailbreak success rates while also holding the potential for defending against jailbreaks.
Forward citations
Cited by 11 Pith papers
-
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
A trigger-tag watermark embedded by fine-tuning lets modified LLMs mark their own phishing outputs for cheap detection.
-
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
A systemization of LLM jailbreak security that adds linked taxonomies, an evaluation platform, and JailbreakDB, while its main attack–defense comparison results remain deferred.
-
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
A small set of attention heads in LVLMs flags malicious prompts during the first token; a logistic-regression detector built on them reduces jailbreak success to 1-5%.
-
VLSBench: Unveiling Visual Leakage in Multimodal Safety
The paper shows existing multimodal safety benchmarks leak harmful image content into text queries (VSIL), and introduces VLSBench, a 2.2k-pair leakless benchmark on which textual alignment fails and multimodal alignm...
-
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
An automated agent can jailbreak GPT-4o and other vision-language models using only individually safe images and benign-sounding prompts, escalating responses to harmful content.
-
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
Near-perfect jailbreak defenses for vision-language models are mostly over-refusal, and the two standard ways of scoring jailbreaks agree only at chance level.
-
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
ArrAttack fine-tunes a judge on the SmoothLLM defense, uses it to filter rewriting-attack data, and trains a generator that produces jailbreak prompts transferring across defenses.
-
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
SAIF is a proposed framework that generates multimodal test prompts from a risk taxonomy, jailbreak tricks, and prompt styles to evaluate generative AI risks in the public sector.
-
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense
Layer-AdvPatcher edits 'toxic' transformer layers using self-generated harmful examples to block jailbreaks, but its reported attack-success rates worsen on several benchmarks.
-
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
A survey that taxonomizes multimodal jailbreak attacks and defenses into four lifecycle levels (input, encoder, generator, output) across Any-to-Text, Any-to-Vision, and Any-to-Any generative models.
-
When Data Manipulation Meets Attack Goals: An In-depth Survey of Attacks for VLMs
A survey that classifies VLM attacks by goal and data manipulation strategy, and reviews defenses and metrics.
Discussion (0). Continue with ORCID to comment.