REVIEW 4 cited by
Perception-guided Jailbreak against Text-to-Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ.
Forward citations
Cited by 4 Pith papers
-
FENCE: A Financial and Multimodal Jailbreak Detection Dataset
FENCE is a new 10k-sample, bilingual, finance-focused multimodal dataset that both exposes VLM jailbreak vulnerabilities and trains small guard models to reject harmful queries with ~99% accuracy.
-
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language Models
A black-box jailbreak method distributes harmful semantics across text and image inputs and uses heuristic search to induce multimodal LLMs to answer harmful queries.
-
An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)
Applying a single-turn 'crescendo' prompt that pretends earlier images were already generated raises DALL-E 3's harmful image rate from 1.3% to 18.5%, close to the uncensored Flux Schnell baseline.
-
Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey
A survey that taxonomizes multimodal jailbreak attacks and defenses into four lifecycle levels (input, encoder, generator, output) across Any-to-Text, Any-to-Vision, and Any-to-Any generative models.
Discussion (0). Continue with ORCID to comment.