Pith. sign in

REVIEW 46 cited by

Red-Teaming the Stable Diffusion Safety Filter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04610 v5 pith:IEGNBUD7 submitted 2022-10-03 cs.AI cs.CRcs.CVcs.CYcs.LG

classification cs.AIcs.CRcs.CVcs.CYcs.LG
keywords filtersafetycontentdiffusionpreventstableaimsdisturbing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Stable Diffusion is a recent open-source image generation model comparable to proprietary models such as DALLE, Imagen, or Parti. Stable Diffusion comes with a safety filter that aims to prevent generating explicit images. Unfortunately, the filter is obfuscated and poorly documented. This makes it hard for users to prevent misuse in their applications, and to understand the filter's limitations and improve it. We first show that it is easy to generate disturbing content that bypasses the safety filter. We then reverse-engineer the filter and find that while it aims to prevent sexual content, it ignores violence, gore, and other similarly disturbing content. Based on our analysis, we argue safety measures in future model releases should strive to be fully open and properly documented to stimulate security contributions from the community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 46 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

    cs.LG 2026-08 conditional novelty 7.0 of 10

    MOON applies spectral-nuclear-norm geometry to multi-objective gradient manipulation and uses polar-factor updates, with O(T^-1/2) deterministic and O(T^-1/4) stochastic convergence to Pareto stationarity.

  2. Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency

    cs.CR 2025-01 conditional novelty 7.0 of 10

    Shuffling words and image patches in harmful prompts bypasses safety mechanisms of several commercial and open-source multimodal models, and a black-box search over shuffles raises attack success rates substantially.

  3. TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

    cs.CR 2026-07 conditional novelty 6.5 of 10

    Safety alignment fails to transfer from text replies to text-in-image, and TYPO’s dual-channel combinatorial search jailbreaks four commercial image models at >90% ASR for ~$0.04.

  4. PEAK: Precise and Persistent Concept Erasure via k-Sparse Autoencoders

    cs.CV 2026-08 conditional novelty 6.0 of 10

    PEAK erases concepts from diffusion models by training a k-sparse autoencoder to localize target features, then fine-tuning the model to suppress those features while preserving all others.

  5. Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    MIND learns a 'defense profile' of a T2I model from fine-grained feedback, then uses it to guide an evolutionary search, achieving 95.62% ASR across six defenses and 91.58% on Wan-2.5.

  6. Evaluating Intellectual Property Guardrails of Generative Image Models: A Technical Report

    cs.CV 2026-07 conditional novelty 6.0 of 10

    All 14 tested text-to-image models readily generate recognizable IP; private models refuse at highly uneven rates, with commercial logos refused least and generated most.

  7. UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

    cs.CV 2026-02 conditional novelty 6.0 of 10

    UnHype generates concept-specific LoRA unlearning weights on the fly from CLIP text embeddings by training a hypernetwork to follow the gradient of an unlearning loss, enabling single- and multi-concept erasure in dif...

  8. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

  9. A Unified Framework for Diffusion Model Unlearning with f-Divergence

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Diffusion model unlearning is generalized from KL/MSE to any f-divergence, with closed-form Hellinger and chi-square losses and a variational min-max form.

  10. LoReUn: Data Itself Implicitly Provides Cues to Improve Machine Unlearning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    LoReUn, a plug-in loss-based reweighting strategy, improves approximate machine unlearning by focusing updates on hard-to-forget low-loss data points.

  11. Customize Multi-modal RAI Guardrails with Precedent-based predictions

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Conditioning a multimodal guardrail on retrieved 'precedent' reasoning traces, rather than static policy definitions, improves few-shot and novel-policy content-moderation F1 scores on UnsafeBench.

  12. PLA: Prompt Learning Attack against Text-to-Image Generative Models

    cs.CR 2025-07 conditional novelty 6.0 of 10

    PLA trains adversarial prompts with a zero-order gradient method and multimodal CLIP losses to bypass safety filters and post-hoc checkers in black-box text-to-image models, outperforming earlier word-substitution attacks.

  13. "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products

    cs.HC 2025-06 conditional novelty 6.0 of 10

    GAI tools' content moderation policies are comprehensive in scope but thin on user reporting and appeals, and Reddit users report frequent frustration with opaque moderation decisions.

  14. GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A reinforcement-learning red-team LLM generates stealthy prompts that bypass text-to-image safety filters and produce toxic images, with reported transfer success against commercial APIs.

  15. SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing

    cs.CV 2025-06 reject novelty 6.0 of 10

    SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.

  16. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  17. Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    RPG-RT iteratively fine-tunes an LLM with rule-based preferences from a decoupled CLIP scoring model, letting it rewrite prompts that bypass unknown safety defenses in black-box text-to-image systems.

  18. TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact Analysis

    cs.CR 2025-05 conditional novelty 6.0 of 10

    TokenProber bypasses five NSFW safety checkers on three text-to-image models by separately preserving dirty words and reducing the influence of non-dirty discrepant words.

  19. Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

    cs.CR 2025-04 conditional novelty 6.0 of 10

    Backdoor triggers injected into text-to-image diffusion models survive state-of-the-art concept erasure, restoring erased identities and explicit content with high success rates.

  20. Predictive Red Teaming: Breaking Policies Without Breaking Robots

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A generative image editing plus anomaly detection pipeline predicts a visuomotor policy's success-rate degradation across off-nominal environmental factors, with an average prediction error below 0.19 in hardware trials.

  21. SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

    cs.LG 2025-01 conditional novelty 6.0 of 10

    SAeUron removes concepts from text-to-image diffusion models by ablating concept-specific sparse autoencoder features during inference, achieving state-of-the-art unlearning on UnlearnCanvas and I2P without weight updates.

  22. T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    T2ISafety is a large annotated benchmark plus a fine-tuned MLLM evaluator (ImageGuard) for measuring toxicity, privacy, and fairness in text-to-image models.

  23. ACE: Anti-Editing Concept Erasure in Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.

  24. SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.

  25. Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A bilevel training procedure that simultaneously restores a pruned diffusion model's quality and suppresses targeted concepts beats sequential fine-tuning followed by unlearning.

  26. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  27. Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Summarizing LLM-obfuscated text-to-image prompts before classification improves content-moderation F1 scores on the new ATTIP dataset.

  28. Moderating the Generalization of Score-based Generative Model

    cs.LG 2024-12 conditional novelty 6.0 of 10

    MSGM is a score-adjustment unlearning method for score-based generative models that suppresses targeted content generation without full retraining.

  29. Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.

  30. Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Typography inserted into input images can manipulate CLIP-guided image generation models to produce harmful or biased content, and existing text-focused defenses do not catch it.

  31. Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Prompt-Noise Optimization jointly tunes the prompt embedding and diffusion noise at inference time to suppress unsafe images while keeping outputs close to the prompt.

  32. Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion Models

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Fine-tuning text-to-image diffusion models on benign data can reactivate suppressed unsafe concepts, and training the task adapter separately from a frozen safety LoRA prevents this.

  33. ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A new guidance reweighting, derived from a contrastive loss, makes negative prompting in diffusion models remove unwanted concepts with less quality loss than standard negated CFG.

  34. DECAF: De-Clustering for Adaptive Representational Unlearning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    DECAF is a forget-only unlearning method that adds input noise, suppresses the forget-class probability, and diversifies outputs, achieving 0.10% forget accuracy and 79.4% retain accuracy on CIFAR-10/ResNet-18 while d...

  35. Introspective Attention Modulation for Safe Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Inference-time attention modulation suppresses unsafe content in diffusion-transformer T2I models without retraining and beats concept-erasure baselines in the paper's benchmarks.

  36. GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention

    cs.CR 2025-07 conditional novelty 5.0 of 10

    GIFT immunizes diffusion models against malicious fine-tuning by combining loss maximization and representation noising, preserving safe concept generation.

  37. Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A new evaluation tool uses (vision-)language model world knowledge to rank nearby concepts and craft adversarial prompts, showing that diffusion unlearning is incomplete and that semantic similarity correlates with co...

  38. The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI

    cs.CR 2025-01 conditional novelty 5.0 of 10

    Security's hard-won transparency practices, from Kerckhoffs' principle to vulnerability disclosure, form three transferable themes for AI transparency efforts.

  39. DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.

  40. Concept Replacer: Replacing Sensitive Concepts in Diffusion Models via Precision Localization

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A method to precisely localize and replace concepts in generated images by combining a few-shot attention-based concept localizer with masked dual-prompt cross-attention during denoising.

  41. Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A step-by-step multimodal 'chain of attack' improves the transferability of targeted adversarial images against open vision-language models, with a new LLM-judged success metric.

  42. MapRoute++: Surrogate-Guided Semantic Routing for Visual Concept Unlearning

    cs.CV 2026-08 conditional novelty 4.0 of 10

    A concept-erasure method built on MapRoute raises the official ERR score from 0.600 to 0.721 on the Genµ2.0 benchmark, but it does not beat MapRoute on object, animal, or action categories.

  43. A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination

    cs.CR 2026-08 conditional novelty 4.0 of 10

    A new jailbreak framework, HACA, combines atomic text and image attack strategies selected by a cross-modal planner and generates attacks with LLMs and text-to-image models, reaching 95.48% average attack success acro...

  44. Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A closed-form concept erasure method that enforces zero alignment residual in the optimization objective and applies updates progressively across layers to better preserve generation quality.

  45. Memory Enhanced Fractional-Order Dung Beetle Optimization for Photovoltaic Parameter Identification

    cs.NE 2025-08 reject novelty 3.0 of 10

    The claimed MFO-DBO algorithm and its CEC2017/PV results are absent from the manuscript, which instead contains an unrelated prompt-stealing attack paper.

  46. Text-to-Image Synthesis: A Decade Survey

    cs.CV 2024-11 conditional novelty 1.0 of 10

    A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.

Pith tools