Pith. sign in

REVIEW 1 cited by

Perception-guided Jailbreak against Text-to-Image Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.10848 v4 pith:FNJS74UC submitted 2024-08-20 cs.CV

classification cs.CV
keywords jailbreakmodelshumanmethodperception-guidedpromptsproposesemantics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW) images. In this paper, inspired by the observation that texts with different semantics can lead to similar human perceptions, we propose an LLM-driven perception-guided jailbreak method, termed PGJ. It is a black-box jailbreak method that requires no specific T2I model (model-free) and generates highly natural attack prompts. Specifically, we propose identifying a safe phrase that is similar in human perception yet inconsistent in text semantics with the target unsafe word and using it as a substitution. The experiments conducted on six open-source models and commercial online services with thousands of prompts have verified the effectiveness of PGJ.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FENCE: A Financial and Multimodal Jailbreak Detection Dataset

    cs.CL 2026-02 conditional novelty 6.0 of 10

    FENCE is a new 10k-sample, bilingual, finance-focused multimodal dataset that both exposes VLM jailbreak vulnerabilities and trains small guard models to reject harmful queries with ~99% accuracy.

Pith tools