Pith. sign in

REVIEW 5 cited by

Erasing Concepts from Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07345 v3 pith:JSQCD7AR submitted 2023-03-13 cs.CV

classification cs.CV
keywords diffusionmodelconceptserasingconductexplicitmethodprevious
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Motivated by recent advancements in text-to-image diffusion, we study erasure of specific concepts from the model's weights. While Stable Diffusion has shown promise in producing explicit or realistic artwork, it has raised concerns regarding its potential for misuse. We propose a fine-tuning method that can erase a visual concept from a pre-trained diffusion model, given only the name of the style and using negative guidance as a teacher. We benchmark our method against previous approaches that remove sexually explicit content and demonstrate its effectiveness, performing on par with Safe Latent Diffusion and censored training. To evaluate artistic style removal, we conduct experiments erasing five modern artists from the network and conduct a user study to assess the human perception of the removed styles. Unlike previous methods, our approach can remove concepts from a diffusion model permanently rather than modifying the output at the inference time, so it cannot be circumvented even if a user has access to model weights. Our code, data, and results are available at https://erasing.baulab.info/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Cross-modal unlearning transfer in vision-language models is asymmetric, architecture-dependent, and shallow under typographic attacks; influence-guided block selection reduces the measured gap.

  2. GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

    cs.LG 2026-01 reject novelty 6.0 of 10

    GUDA approximates leave-one-group-out counterfactual models with unlearning and ranks group influence by ELBO differences.

  3. Few to Big: Prototype Expansion Network via Diffusion Learner for Point Cloud Few-shot Semantic Segmentation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    PENet expands few-shot point cloud prototypes by combining a standard supervised encoder with a repurposed diffusion-model encoder and reports SOTA mIoU on S3DIS and ScanNet.

  4. Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Water4MU tunes an invisible watermark on data so that machine unlearning algorithms can remove requested images more effectively, beating prior methods on 'challenging forgets'.

  5. TRACE: Trajectory-Constrained Concept Erasure in Diffusion Models

    cs.CV 2025-05 reject novelty 2.0 of 10

    TRACE combines a closed-form cross-attention nullification with a late-timestep fine-tuning loss to erase concepts from diffusion models, claiming better erasure and fidelity than published baselines.

Pith tools