Pith. sign in

REVIEW 5 cited by

Towards Safe Self-Distillation of Internet-Scale Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.05977 v1 pith:3QLVNAAB submitted 2023-07-12 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelscontentdiffusionmethodconceptgenerationharmfulimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale image generation models, with impressive quality made possible by the vast amount of data available on the Internet, raise social concerns that these models may generate harmful or copyrighted content. The biases and harmfulness arise throughout the entire training process and are hard to completely remove, which have become significant hurdles to the safe deployment of these models. In this paper, we propose a method called SDD to prevent problematic content generation in text-to-image diffusion models. We self-distill the diffusion model to guide the noise estimate conditioned on the target removal concept to match the unconditional one. Compared to the previous methods, our method eliminates a much greater proportion of harmful content from the generated images without degrading the overall image quality. Furthermore, our method allows the removal of multiple concepts at once, whereas previous works are limited to removing a single concept at a time.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LU-500: A Logo Benchmark for Concept Unlearning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new 500-company benchmark shows current concept-erasure methods cannot remove small logos from generated images without also changing unrelated content.

  2. SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing

    cs.CV 2025-06 reject novelty 6.0 of 10

    SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.

  3. Model Immunization from a Condition Number Perspective

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new regularizer increases the condition number of the linear-probing Hessian on harmful tasks, making gradient-descent fine-tuning slower, but the theoretical analysis contains a false claim.

  4. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  5. SecureT2I: No More Unauthorized Manipulation on AI Generated Images from Prompts

    cs.CR 2025-07 reject novelty 4.0 of 10

    A diffusion editing model is fine-tuned with a blur target for forbidden images and the original output for permitted images, claiming selective suppression of unauthorized edits.

Pith tools