Pith. sign in

REVIEW 15 cited by

Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15234 v3 pith:3FFDZBZK submitted 2024-05-24 cs.CV cs.CR

classification cs.CVcs.CR
keywords unlearningadvunlearnconceptrobustnessrobustadversarialerasingerasure
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adversarial prompt attacks, which can prompt DMs post-unlearning to regenerate undesired images containing concepts (such as nudity) meant to be erased. This work aims to enhance the robustness of concept erasing by integrating the principle of adversarial training (AT) into machine unlearning, resulting in the robust unlearning framework referred to as AdvUnlearn. However, achieving this effectively and efficiently is highly nontrivial. First, we find that a straightforward implementation of AT compromises DMs' image generation quality post-unlearning. To address this, we develop a utility-retaining regularization on an additional retain set, optimizing the trade-off between concept erasure robustness and model utility in AdvUnlearn. Moreover, we identify the text encoder as a more suitable module for robustification compared to UNet, ensuring unlearning effectiveness. And the acquired text encoder can serve as a plug-and-play robust unlearner for various DM types. Empirically, we perform extensive experiments to demonstrate the robustness advantage of AdvUnlearn across various DM unlearning scenarios, including the erasure of nudity, objects, and style concepts. In addition to robustness, AdvUnlearn also achieves a balanced tradeoff with model utility. To our knowledge, this is the first work to systematically explore robust DM unlearning through AT, setting it apart from existing methods that overlook robustness in concept erasing. Codes are available at: https://github.com/OPTML-Group/AdvUnlearn

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models

    cs.CV 2025-09 conditional novelty 7.0 of 10

    SuMa erases narrow concepts from text-to-image models by mapping the concept's token subspace onto a nearby reference subspace, achieving robustness against adversarial attacks with image quality close to standard era...

  2. Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing

    cs.CV 2024-11 conditional novelty 7.0 of 10

    FaceLock perturbs portraits so diffusion-based edits destroy face-recognition similarity, and it evaluates success with the same face model that it attacks.

  3. Minimalist Concept Erasure in Generative Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A final-output-only loss with learned neuron masks erases concepts from flow-based image generators more robustly than per-step fine-tuning methods.

  4. Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CPE uses nonlinear residual attention gates with anchoring and adversarial training to erase target concepts from text-to-image diffusion models while preserving remaining concepts better than prior fine-tuning methods.

  5. SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing

    cs.CV 2025-06 reject novelty 6.0 of 10

    SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.

  6. Knowledge Swapping via Learning and Unlearning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Learning new knowledge first and then forgetting selected classes outperforms the reverse order for a pretrained vision model, on classification, segmentation, and detection.

  7. ACE: Anti-Editing Concept Erasure in Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.

  8. Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A bilevel training procedure that simultaneously restores a pruned diffusion model's quality and suppresses targeted concepts beats sequential fine-tuning followed by unlearning.

  9. TraSCE: Trajectory Steering for Concept Erasure

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TraSCE steers diffusion trajectories with a modified negative-prompt formulation and a Gaussian loss to erase concepts at inference time without training or weight updates.

  10. Memories of Forgotten Concepts

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Erased concepts in text-to-image diffusion models can still be generated from high-likelihood latent seeds recovered by diffusion inversion, across nine ablation methods and six concepts.

  11. Opt-In Art: Learning Art Styles Only from Few Examples

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A diffusion model pretrained exclusively on photographs can learn a painter's style from just a handful of examples, matching the style fidelity of models pretrained on large art-containing datasets.

  12. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  13. Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification

    cs.LG 2024-12 reject novelty 5.0 of 10

    A single optimized input perturbation can make a fixed image classifier misclassify targeted classes, mimicking unlearning without any weight update.

  14. MUNBa: Machine Unlearning via Nash Bargaining

    cs.CV 2024-11 conditional novelty 5.0 of 10

    MUNBa is a machine unlearning method that uses Nash bargaining to balance forgetting and preservation gradients, improving unlearning quality, generalization, and robustness in image classification and generation.

  15. FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models

    cs.CV 2024-12 conditional novelty 4.0 of 10

    FameBias linearly combines a famous person's embedding with a trigger word's embedding to make text-to-image models generate that person, reaching 53% bias success without training.

Pith tools