Pith. sign in

REVIEW 7 cited by

Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12383 v2 pith:F3CPOODL submitted 2024-07-17 cs.CV

classification cs.CV
keywords erasureconceptsreceefficientabilityembeddingsgenerationinappropriate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image models encounter safety issues, including concerns related to copyright and Not-Safe-For-Work (NSFW) content. Despite several methods have been proposed for erasing inappropriate concepts from diffusion models, they often exhibit incomplete erasure, consume a lot of computing resources, and inadvertently damage generation ability. In this work, we introduce Reliable and Efficient Concept Erasure (RECE), a novel approach that modifies the model in 3 seconds without necessitating additional fine-tuning. Specifically, RECE efficiently leverages a closed-form solution to derive new target embeddings, which are capable of regenerating erased concepts within the unlearned model. To mitigate inappropriate content potentially represented by derived embeddings, RECE further aligns them with harmless concepts in cross-attention layers. The derivation and erasure of new representation embeddings are conducted iteratively to achieve a thorough erasure of inappropriate concepts. Besides, to preserve the model's generation ability, RECE introduces an additional regularization term during the derivation process, resulting in minimizing the impact on unrelated concepts during the erasure process. All the processes above are in closed-form, guaranteeing extremely efficient erasure in only 3 seconds. Benchmarking against previous approaches, our method achieves more efficient and thorough erasure with minor damage to original generation ability and demonstrates enhanced robustness against red-teaming tools. Code is available at \url{https://github.com/CharlesGong12/RECE}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CPE uses nonlinear residual attention gates with anchoring and adversarial training to erase target concepts from text-to-image diffusion models while preserving remaining concepts better than prior fine-tuning methods.

  2. ACE: Anti-Editing Concept Erasure in Text-to-Image Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    ACE trains a LoRA adapter on both conditional and unconditional noise predictions so that erased concepts are suppressed during both generation and text-guided editing.

  3. SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SafeCFG adapts classifier-free guidance with a learned feature controller so that clean prompts generate normally while harmful prompts are pushed away from unsafe content.

  4. TraSCE: Trajectory Steering for Concept Erasure

    cs.CV 2024-12 conditional novelty 6.0 of 10

    TraSCE steers diffusion trajectories with a modified negative-prompt formulation and a Gaussian loss to erase concepts at inference time without training or weight updates.

  5. Moderating the Generalization of Score-based Generative Model

    cs.LG 2024-12 conditional novelty 6.0 of 10

    MSGM is a score-adjustment unlearning method for score-based generative models that suppresses targeted content generation without full retraining.

  6. DuMo: Dual Encoder Modulation Network for Precise Concept Erasure

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DuMo erases target concepts from text-to-image models by adding a frozen-backbone skip-connection eraser with learned timestep and layer modulation, reporting the best trade-off on three concept erasure benchmarks.

  7. FameBias: Embedding Manipulation Bias Attack in Text-to-Image Models

    cs.CV 2024-12 conditional novelty 4.0 of 10

    FameBias linearly combines a famous person's embedding with a trigger word's embedding to make text-to-image models generate that person, reaching 53% bias success without training.

Pith tools