Pith. sign in

REVIEW 1 cited by

DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.11071 v2 pith:CJ2JNA2F submitted 2024-08-18 cs.CR cs.AIcs.CV

classification cs.CRcs.AIcs.CV
keywords attackmodeldiffzoomechanismssafetyteamingattacksblack-box
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current text-to-image (T2I) synthesis diffusion models raise misuse concerns, particularly in creating prohibited or not-safe-for-work (NSFW) images. To address this, various safety mechanisms and red teaming attack methods are proposed to enhance or expose the T2I model's capability to generate unsuitable content. However, many red teaming attack methods assume knowledge of the text encoders, limiting their practical usage. In this work, we rethink the case of \textit{purely black-box} attacks without prior knowledge of the T2l model. To overcome the unavailability of gradients and the inability to optimize attacks within a discrete prompt space, we propose DiffZOO which applies Zeroth Order Optimization to procure gradient approximations and harnesses both C-PRV and D-PRV to enhance attack prompts within the discrete prompt domain. We evaluated our method across multiple safety mechanisms of the T2I diffusion model and online servers. Experiments on multiple state-of-the-art safety mechanisms show that DiffZOO attains an 8.5% higher average attack success rate than previous works, hence its promise as a practical red teaming tool for T2l models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ZIUM attacks unlearned diffusion models by optimizing an image-captioning module that turns a target image into a text embedding, then reuses that module zero-shot on unseen images of the same unlearned concept.

Pith tools