REVIEW 15 cited by
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models (DMs) have achieved remarkable success in text-to-image generation, but they also pose safety risks, such as the potential generation of harmful content and copyright violations. The techniques of machine unlearning, also known as concept erasing, have been developed to address these risks. However, these techniques remain vulnerable to adversarial prompt attacks, which can prompt DMs post-unlearning to regenerate undesired images containing concepts (such as nudity) meant to be erased. This work aims to enhance the robustness of concept erasing by integrating the principle of adversarial training (AT) into machine unlearning, resulting in the robust unlearning framework referred to as AdvUnlearn. However, achieving this effectively and efficiently is highly nontrivial. First, we find that a straightforward implementation of AT compromises DMs' image generation quality post-unlearning. To address this, we develop a utility-retaining regularization on an additional retain set, optimizing the trade-off between concept erasure robustness and model utility in AdvUnlearn. Moreover, we identify the text encoder as a more suitable module for robustification compared to UNet, ensuring unlearning effectiveness. And the acquired text encoder can serve as a plug-and-play robust unlearner for various DM types. Empirically, we perform extensive experiments to demonstrate the robustness advantage of AdvUnlearn across various DM unlearning scenarios, including the erasure of nudity, objects, and style concepts. In addition to robustness, AdvUnlearn also achieves a balanced tradeoff with model utility. To our knowledge, this is the first work to systematically explore robust DM unlearning through AT, setting it apart from existing methods that overlook robustness in concept erasing. Codes are available at: https://github.com/OPTML-Group/AdvUnlearn
Forward citations
Cited by 15 Pith papers
-
SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
SuMa erases narrow concepts from text-to-image models by mapping the concept's token subspace onto a nearby reference subspace, achieving robustness against adversarial attacks with image quality close to standard era...
-
Edit Away and My Face Will not Stay: Personal Biometric Defense against Malicious Generative Editing
FaceLock perturbs portraits so diffusion-based edits destroy face-recognition similarity, and it evaluates success with the same face model that it attacks.
-
Minimalist Concept Erasure in Generative Models
A final-output-only loss with learned neuron masks erases concepts from flow-based image generators more robustly than per-step fine-tuning methods.
-
Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
CPE uses nonlinear residual attention gates with anchoring and adversarial training to erase target concepts from text-to-image diffusion models while preserving remaining concepts better than prior fine-tuning methods.
-
SAGE: Exploring the Boundaries of Unsafe Concept Domain with Semantic-Augment Erasing
SAGE erases concepts from diffusion models by optimizing attack prompts against the model's own text encoder and then fine-tuning that encoder with a global-local retention loss.
-
Knowledge Swapping via Learning and Unlearning
Learning new knowledge first and then forgetting selected classes outperforms the reverse order for a pretrained vision model, on classification, segmentation, and detection.
-
ACE: Anti-Editing Concept Erasure in Text-to-Image Models
ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.
-
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
A bilevel training procedure that simultaneously restores a pruned diffusion model's quality and suppresses targeted concepts beats sequential fine-tuning followed by unlearning.
-
TraSCE: Trajectory Steering for Concept Erasure
TraSCE steers diffusion trajectories with a modified negative-prompt formulation and a Gaussian loss to erase concepts at inference time without training or weight updates.
-
Memories of Forgotten Concepts
Erased concepts in text-to-image diffusion models can still be generated from high-likelihood latent seeds recovered by diffusion inversion, across nine ablation methods and six concepts.
-
Opt-In Art: Learning Art Styles Only from Few Examples
A diffusion model pretrained exclusively on photographs can learn a painter's style from just a handful of examples, matching the style fidelity of models pretrained on large art-containing datasets.
-
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.
-
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
A single optimized input perturbation can make a fixed image classifier misclassify targeted classes, mimicking unlearning without any weight update.
-
MUNBa: Machine Unlearning via Nash Bargaining
MUNBa is a machine unlearning method that uses Nash bargaining to balance forgetting and preservation gradients, improving unlearning quality, generalization, and robustness in image classification and generation.
-
FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models
FameBias linearly combines a famous person's embedding with a trigger word's embedding to make text-to-image models generate that person, reaching 53% bias success without training.
Discussion (0). Continue with ORCID to comment.