Pith. sign in

REVIEW 15 cited by

Erasing Undesirable Influence in Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.05779 v4 pith:UZLH2CCU submitted 2024-01-11 cs.CV

classification cs.CV
keywords diffusionmodelswhilealgorithmdataeffectiveerasediffmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models are highly effective at generating high-quality images but pose risks, such as the unintentional generation of NSFW (not safe for work) content. Although various techniques have been proposed to mitigate unwanted influences in diffusion models while preserving overall performance, achieving a balance between these goals remains challenging. In this work, we introduce EraseDiff, an algorithm designed to preserve the utility of the diffusion model on retained data while removing the unwanted information associated with the data to be forgotten. Our approach formulates this task as a constrained optimization problem using the value function, resulting in a natural first-order algorithm for solving the optimization problem. By altering the generative process to deviate away from the ground-truth denoising trajectory, we update parameters for preservation while controlling constraint reduction to ensure effective erasure, striking an optimal trade-off. Extensive experiments and thorough comparisons with state-of-the-art algorithms demonstrate that EraseDiff effectively preserves the model's utility, efficacy, and efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

    cs.LG 2026-08 conditional novelty 7.0 of 10

    MOON applies spectral-nuclear-norm geometry to multi-objective gradient manipulation and uses polar-factor updates, with O(T^-1/2) deterministic and O(T^-1/4) stochastic convergence to Pareto stationarity.

  2. UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

    cs.CV 2026-02 conditional novelty 6.0 of 10

    UnHype generates concept-specific LoRA unlearning weights on the fly from CLIP text embeddings by training a hypernetwork to follow the gradient of an unlearning loss, enabling single- and multi-concept erasure in dif...

  3. SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A supervised sparse autoencoder binds each concept to a single neuron, letting Stable Diffusion erase a concept by steering one latent.

  4. LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

    cs.CR 2025-07 conditional novelty 6.0 of 10

    A data-free editing framework that realigns LoRA weight subspaces via adversarial optimization and semantic augmentation to suppress malicious generations.

  5. Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A diffusion model unlearning method that automatically chooses a closely related, non-synonym target concept for each erased concept, reducing side effects on other concepts while keeping erasure effective.

  6. SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

    cs.LG 2025-01 conditional novelty 6.0 of 10

    SAeUron removes concepts from text-to-image diffusion models by ablating concept-specific sparse autoencoder features during inference, achieving state-of-the-art unlearning on UnlearnCanvas and I2P without weight updates.

  7. SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.

  8. Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A bilevel training procedure that simultaneously restores a pruned diffusion model's quality and suppresses targeted concepts beats sequential fine-tuning followed by unlearning.

  9. Moderating the Generalization of Score-based Generative Model

    cs.LG 2024-12 conditional novelty 6.0 of 10

    MSGM is a score-adjustment unlearning method for score-based generative models that suppresses targeted content generation without full retraining.

  10. Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.

  11. Memories of Forgotten Concepts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Erased concepts in text-to-image diffusion models can still be generated from high-likelihood latent seeds recovered by diffusion inversion, across nine ablation methods and six concepts.

  12. ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A new guidance reweighting, derived from a contrastive loss, makes negative prompting in diffusion models remove unwanted concepts with less quality loss than standard negated CFG.

  13. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  14. Few-Shot Concept Unlearning with Low Rank Adaptation

    cs.LG 2025-05 reject novelty 5.0 of 10

    The authors combine few-shot unlearning with low-rank adaptation on the CLIP text encoder to erase concepts from Stable Diffusion v2 in under a minute, reporting low forget-CLIP scores and detection rates on three concepts.

  15. Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A survey that taxonomizes multimodal jailbreak attacks and defenses into four lifecycle levels (input, encoder, generator, output) across Any-to-Text, Any-to-Vision, and Any-to-Any generative models.

Pith tools