Pith. sign in

REVIEW 6 cited by

Circumventing Concept Erasure Methods For Text-to-Image Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01508 v2 pith:VHSEF624 submitted 2023-08-03 cs.LG cs.CRcs.CV

classification cs.LGcs.CRcs.CV
keywords methodsmodelsconceptsconcepterasuretext-to-imagegenerativeimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image generative models can produce photo-realistic images for an extremely broad range of concepts, and their usage has proliferated widely among the general public. On the flip side, these models have numerous drawbacks, including their potential to generate images featuring sexually explicit content, mirror artistic styles without permission, or even hallucinate (or deepfake) the likenesses of celebrities. Consequently, various methods have been proposed in order to "erase" sensitive concepts from text-to-image models. In this work, we examine five recently proposed concept erasure methods, and show that targeted concepts are not fully excised from any of these methods. Specifically, we leverage the existence of special learned word embeddings that can retrieve "erased" concepts from the sanitized models with no alterations to their weights. Our results highlight the brittleness of post hoc concept erasure methods, and call into question their use in the algorithmic toolkit for AI safety.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Diffusion-grounded erase/retain retrieval plus retain-orthogonal value projection and trigger-guided subspace expansion erases concepts more robustly than prior CETs while keeping FID/CLIP near the unedited model.

  2. Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    ConceptAgent is a black-box multi-agent system that awakens erased concepts in diffusion models by initializing denoising trajectories from surrogate-guided noisy states.

  3. When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

    cs.CR 2026-04 unverdicted novelty 6.0 of 10

    TwoHamsters benchmark shows T2I models like FLUX generate unsafe multi-concept images at 99.52% rate while defenses like LLaVA-Guard achieve only 41.06% recall.

  4. When Safe Concepts Become Unsafe: Multi-Concept Compositional Vulnerabilities in Text-to-Image Models

    cs.CR 2026-04 unverdicted novelty 6.0 of 10

    Combining multiple safe visual concepts in one prompt can make text-to-image models produce harmful images, and stronger instruction-following models fail more often.

  5. Co-occurring associated retained concepts in Diffusion Unlearning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Defines CARE score and proposes ReCARE framework to preserve co-occurring benign concepts during targeted unlearning in diffusion models.

  6. ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline

    cs.CR 2026-02 conditional novelty 4.0 of 10

    ReVision uses a vision-language model's bounding box to gate attention-based image editing, suppressing unsafe concepts while better preserving benign background in multi-concept scenes.

Pith tools