REVIEW 3 cited by
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale pre-trained generative models are taking the world by storm, due to their abilities in generating creative content. Meanwhile, safeguards for these generative models are developed, to protect users' rights and safety, most of which are designed for large language models. Existing methods primarily focus on jailbreak and adversarial attacks, which mainly evaluate the model's safety under malicious prompts. Recent work found that manually crafted safe prompts can unintentionally trigger unsafe generations. To further systematically evaluate the safety risks of text-to-image models, we propose a novel Automatic Red-Teaming framework, ART. Our method leverages both vision language model and large language model to establish a connection between unsafe generations and their prompts, thereby more efficiently identifying the model's vulnerabilities. With our comprehensive experiments, we reveal the toxicity of the popular open-source text-to-image models. The experiments also validate the effectiveness, adaptability, and great diversity of ART. Additionally, we introduce three large-scale red-teaming datasets for studying the safety risks associated with text-to-image models. Datasets and models can be found in https://github.com/GuanlinLee/ART.
Forward citations
Cited by 3 Pith papers
-
GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models
A reinforcement-learning red-team LLM generates stealthy prompts that bypass text-to-image safety filters and produce toxic images, with reported transfer success against commercial APIs.
-
Any-Resolution AI-Generated Image Detection by Spectral Learning
SPAI uses spectral reconstruction similarity from a frozen masked-frequency ViT plus attention pooling to reach 91.0 average AUC for AI-generated image detection across 13 unseen generators.
-
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models
An automated agent can jailbreak GPT-4o and other vision-language models using only individually safe images and benign-sounding prompts, escalating responses to harmful content.
Discussion (0). Continue with ORCID to comment.