Zero-query jailbreak attacks on text-to-image systems that exploit filter-generator discrepancy reach 29-33% average success and beat baselines on six pipelines and GPT-image-2.
Discovering the Hidden Vocabulary of DALLE-2
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We discover that DALLE-2 seems to have a hidden vocabulary that can be used to generate images with absurd prompts. For example, it seems that \texttt{Apoploe vesrreaitais} means birds and \texttt{Contarra ccetnxniams luryca tanniounons} (sometimes) means bugs or pests. We find that these prompts are often consistent in isolation but also sometimes in combinations. We present our black-box method to discover words that seem random but have some correspondence to visual concepts. This creates important security and interpretability challenges.
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems
Zero-query jailbreak attacks on text-to-image systems that exploit filter-generator discrepancy reach 29-33% average success and beat baselines on six pipelines and GPT-image-2.