{"id":"36faf789-37a2-459b-b950-2dd839de99d9","arxiv_id":"2505.19398","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"This survey classifies concept erasure methods for text-to-image diffusion models along intervention level, optimization strategy, and semantic scope, and reviews the datasets, metrics, and benchmarks used to evaluate them.","lead":"This preprint surveys methods for erasing concepts such as nudity, violence, or artistic styles from text-to-image diffusion models. It organizes dozens of techniques into a three-axis taxonomy and reviews benchmarks and open challenges.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Tables I and II in Section III.C are internally inconsistent, so the claimed trade-off insights are not supported.","rationale":"The reader's weakest assumption identified the qualitative Effiveness/Utility ratings in Tables I and II as the load-bearing element of the survey's comparative contribution. My stress-test confirms this and sharpens it: the problem is not merely the absence of a cited evaluation protocol, but actual internal inconsistency between the two tables. Because Table I and Table II are presented as two orientations of the same comparison, the same intervention-strategy pair must receive the same ratings in both. Several cells do not, including Text Encoder + Loss-based Utility, Text Encoder + Closed-form Utility, Text Encoder + Adversarial Effectiveness/Utility, UNet + Loss-based Effectiveness/Utility, and Cross-Attention + Plug-in Adapters Utility. These contradictions cannot be attributed to differing evaluation standards between papers, because within a single paper the two tables should encode one rating matrix. I considered other reported issues, such as the mis-cited Safety Checker reference [24], the withdrawn paper [74] presented without qualification, and notation slips in equations; these are local and correctable, and they do not undermine the taxonomy itself. The internal inconsistency directly affects the claimed 'actionable insights' and 'clearer methodological comparisons,' which are central to the survey's stated purpose. The right disposition remains CONDITIONAL: the taxonomy and coverage can still be valuable, but Section III.C's comparative conclusions should not be accepted as-is until the tables are reconciled with an explicit, source-based rating protocol or replaced with quantitative per-paper evidence. This does not change the reader's verdict, but it provides a concrete, checkable reason to insist on revision.","tokens_in":27301,"tokens_out":5304,"duration_ms":51578,"concrete_test":"Construct the canonical cross-tabulation of intervention level (Text Encoder, Cross-Attention, UNet) by optimization strategy (Loss-based, Closed-form, Plug-in Adapters, Adversarial). For each of the 12 combinations, extract the Effectiveness and Utility labels from both Table I and Table II and compare them cell by cell. If any cell differs, the two tables assert conflicting ratings for the same combination, and the trade-off narrative built on them is unreliable. Then, for a sample of the mismatched cells, list the papers in Section III that instantiate that combination and look up their reported ESR, FID, and robustness numbers; if no explicit binning rule maps those numbers to Low/Moderate/High, or if the implied binning differs from one of the tables, the qualitative ratings should be removed or replaced with a per-paper quantitative comparison table.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The survey's central claim is that its taxonomy enables 'clearer methodological comparisons' and yields 'actionable insights' about erasure-effectiveness versus utility trade-offs. In Section III.C, those insights rest on Tables I and II, which assign qualitative Effectiveness and Utility ratings to each intervention-level by optimization-strategy combination. These ratings are presented without an evaluation protocol, a cited source, or a mapping from reported metrics such as ESR, FID, and robustness. More importantly, the two tables are meant to be transposed views of the same rating matrix, and they contradict each other. For example: Text Encoder + Loss-based has Utility 'Moderate' in Table I but 'High' in Table II; Text Encoder + Closed-form has Utility 'Moderate' in Table I but 'High' in Table II; Text Encoder + Adversarial has Effectiveness 'High' and Utility 'Low-Moderate' in Table I but 'Moderate' and 'Moderate' in Table II; UNet + Loss-based has Effectiveness 'Moderate' and Utility 'Moderate' in Table I but 'High' and 'Low-Moderate' in Table II; Cross-Attention + Plug-in Adapters has Utility 'Moderate-High' in Table I but 'High' in Table II. These are not subtle gradations; they are the same cells rated differently in two tables that are supposed to summarize the same comparison. Since the tables are the only support for the claimed qualitative trade-offs, the comparative conclusions in Section III.C are not currently reproducible or internally consistent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews concept-erasure methods for text-to-image diffusion models, organizing the literature along three axes: intervention level (text encoder, cross-attention, UNet), optimization strategy (loss-based, closed-form, plug-in adapters, adversarial), and semantic scope (single, multi, compositional). It provides background definitions, a taxonomy diagram, equations for representative methods, qualitative effectiveness/utility ratings in Tables I and II, an overview of datasets and metrics, and discussion of challenges and future directions. The stated contribution is a structured framework enabling 'clearer methodological comparisons' and 'actionable insights' into erasure-utility trade-offs.","tokens_in":27523,"tokens_out":8196,"duration_ms":76532,"significance":"The survey addresses a rapidly growing and practically important area, and its temporal coverage through May 2025 is a genuine asset. The proposed taxonomy is plausible and largely consistent with the cited methods, and the compilation of datasets, metrics, and benchmark frameworks in Section IV is useful for practitioners. The paper does not present new experimental results, so its value rests on the accuracy, completeness, and internal consistency of its qualitative synthesis. If the taxonomy and ratings were made reliable, the survey would serve as a useful entry point for method selection; at present, the central comparative synthesis is undermined by unsupported and internally contradictory ratings in Tables I and II, and by several citation and notation errors that reduce confidence in the survey's accuracy.","major_comments":[{"comment":"These two tables are presented as transposed views of the same rating matrix, but several cells disagree. Text Encoder + Loss-based is rated Utility 'Moderate' in Table I and 'High' in Table II; Text Encoder + Closed-form is 'Moderate' in Table I and 'High' in Table II; Text Encoder + Adversarial is Effectiveness 'High' and Utility 'Low-Moderate' in Table I but 'Moderate' and 'Moderate' in Table II; UNet + Loss-based is 'Moderate' and 'Moderate' in Table I but 'High' and 'Low-Moderate' in Table II; Cross-Attention + Plug-in Adapters is Utility 'Moderate-High' in Table I and 'High' in Table II. Since these tables are the only support for the trade-off conclusions in Section III.C.d, the contradictions are load-bearing. The authors must either harmonize the tables and specify the protocol or source for each rating, including the mapping from reported metrics such as ESR, FID, and robustness to the ordinal labels, or explicitly label the ratings as subjective editorial opinion and temper the claims that rest on them.","section":"Section III.C, Tables I and II"},{"comment":"The paper's central claim is a three-dimensional taxonomy, but Tables I and II only compare intervention level and optimization strategy; the third axis, semantic scope, is absent from the comparative synthesis. Section III.C.c discusses semantic scope in prose without any systematic comparison, so the claimed multidimensional comparison is only partially realized. Please either extend the comparison to include semantic scope or revise the claim to describe a two-dimensional comparison with semantic scope as a descriptive category rather than a compared dimension.","section":"Section III.C and Figure 2"},{"comment":"The update rule for Buster is written as y^{(n+1)} = -y^{(n)} - eta * grad E(y^{(n)}) + epsilon^{(n)}. With eta > 0, this is not a gradient-descent step and is inconsistent with the described objective of minimizing E; as written, the iteration diverges rather than converging. Please correct the equation or clarify the notation so that the description matches the cited method.","section":"Section III.A.1.a, Eq. (7)"},{"comment":"Several cited references do not correspond to the methods they are attached to. Figure 2 lists 'CA [97]' as an explicit-concept method, but reference [97] is the CAT cross-attention paper, while the CA method is reference [37]. Section IV.A identifies the 'Safety Checker' with reference [24] (NudeNet), whereas the Stable Diffusion safety checker is a different component typically associated with the model description in reference [99]. In addition, reference [74] is explicitly marked in the bibliography as withdrawn from ICLR 2025, yet it is presented in Section III.A.2.b without qualification. These inaccuracies undercut the survey's promise of being a reliable comparative guide and should be corrected.","section":"Figure 2, Section IV.A, and reference list"}],"minor_comments":[{"comment":"The figure contains typos: 'Meta-Unlearing' should read 'Meta-Unlearning', 'Safety Checher' should read 'Safety Checker', and 'ErasingAnithing' should read 'EraseAnything'.","section":"Figure 2"},{"comment":"Near Eq. (7), 'Guassian' should be 'Gaussian'.","section":"Section III.A.1.a"},{"comment":"The paper would benefit from a short paragraph describing the inclusion criteria and search protocol used to select the reviewed methods, since the claim of comprehensiveness is central to a survey's usefulness.","section":"Section IV"},{"comment":"Two distinct methods named 'ACE' ([52] and [66]) appear without disambiguation; please add descriptors such as 'Anti-editing Concept Erasure' and 'Attentional Concept Erasure' to avoid confusion.","section":"Figure 2 and Section III.A.2.a"},{"comment":"In Eq. (14), the PromptSlider objective does not explicitly indicate the distribution over timestep t in the expectation; specifying it would improve reproducibility of the presented formulation.","section":"Section III.A.1.e"}],"recommendation":"major_revision","confidential_remarks":"The survey's scope is appropriate for a journal survey, and the taxonomy is potentially useful. The main obstacle is not novelty but the unsupported and mutually inconsistent qualitative ratings in Tables I and II, which currently undermine the paper's central comparative claims. The issue is fixable within the manuscript's scope by removing or substantiating those ratings and by correcting the citation and notation errors; I would be comfortable with acceptance after such a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [Colleague],\n\nThe main thing to know: this is a genuinely useful survey of concept erasure in T2I diffusion models, but the two comparative tables at the heart of its 'actionable insights' contradict each other and should not be used as-is.\n\nWhat's new and good: the three-axis taxonomy (intervention level, optimization structure, semantic scope) is a clean way to organize a fragmented literature. Coverage is up to date through May 2025, including recent plug-in adapters, continual erasure, and preference-aligned methods. The evaluation section is a solid summary of metrics (ESR, FID, CLIP score) and benchmarks (I2P, UnlearnCanvas, EraseBench, etc.). If you want a map of the field, this is a usable one. The self-citations in the introduction are just background examples and are not a problem.\n\nWhere it gets soft: Section III.C's Tables I and II are presented as complementary views of the same comparison, but several cells disagree. For instance, Text Encoder + Loss-based has Utility 'Moderate' in Table I but 'High' in Table II; Text Encoder + Adversarial Training has Effectiveness 'High' and Utility 'Low-Moderate' in Table I but 'Moderate' and 'Moderate' in Table II. Since no evaluation protocol or cited source is given for these qualitative ratings, the comparative conclusions in III.C are not currently supported. That is a load-bearing flaw, not a cosmetic one.\n\nOther issues: the 'Safety Checker' is mis-cited as NudeNet [24] and is classified under 'Text Encoder-Based Similarity Filtering,' although the standard safety checker operates on decoded images rather than text encoder representations. Reference [74] is a withdrawn ICLR submission and is treated as a regular method without noting the withdrawal. There are also typos ('Meta-Unlearing,' 'Guassian') that suggest a rushed final pass.\n\nBottom line: the taxonomy and coverage deserve to be published after revision. The qualitative tables need to be either removed, anchored to a reproducible protocol, or reframed as the authors' own judgment with caveats. The citation and category errors are fixable.\n\nWho this is for: researchers entering concept erasure or practitioners selecting a method family. It deserves a serious referee, with revisions expected.","headline":"Useful survey with a solid taxonomy, but the central comparison tables are internally inconsistent and need major fixing.","tokens_in":28055,"tokens_out":3822,"would_cite":false,"duration_ms":32862,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes a three-dimensional taxonomy—intervention level, optimization strategy, and semantic scope—that organizes concept erasure methods in text-to-image diffusion models so practitioners can compare trade-offs between…","keywords":["concept erasure","text-to-image diffusion","machine unlearning","safety alignment","NSFW content","cross-attention","closed-form erasure","plug-in adapters"],"falsifier":"A standardized evaluation that runs one representative method per taxonomy cell on the same prompts, metrics, and random seeds would falsify the survey's trade-off claims if the measured Effectiveness and Utility rankings contradict the ratings in Tables I and II (for example, if a text-encoder-level loss-based method outperformed every UNet-level method in both effectiveness and utility).","tokens_in":27044,"feed_emoji":"🛡️","tokens_out":4346,"duration_ms":37641,"temperature":0.7,"pith_summary":"This survey paper claims that the scattered literature on concept erasure in text-to-image diffusion models can be organized along three independent axes: where in the model the intervention happens (text encoder, cross-attention, or UNet), how the suppression is learned (loss-based optimization, closed-form projections, plug-in adapters, or adversarial training), and what semantic scope the target has (a single explicit concept, multiple concepts, or the combination of individually benign concepts). It argues that this three-dimensional framing makes methodological comparisons clearer and exposes a recurring trade-off between erasure strength and model utility: deeper interventions suppress more reliably but cost more in fidelity and complexity. The survey further collects the evaluation metrics, datasets, and benchmarks used to measure erasure success and utility, and identifies open problems such as conceptual entanglement, reactivation attacks, and adaptation to new model architectures. A careful reader would care because the taxonomy is intended as a practical map for choosing or designing an erasure method for a specific safety requirement.","feed_headline":"Three axes map every concept-erasure method","feed_subtitle":"Survey maps text-to-image erasure by where, how, and what it targets, exposing a strength-versus-utility trade-off.","key_machinery":"The load-bearing object is the taxonomy itself, presented in Figure 2 as a three-axis scheme. The first axis, intervention level, splits methods into text-encoder-level (prompt embedding modifications), cross-attention-level (attention-map or key/value manipulation), and UNet-level (feature masking, pruning, loss-driven). The second axis, optimization structure, distinguishes loss-based objectives, closed-form projections (e.g., least-squares updates of attention projections), plug-in adapters (LoRA and gated modules), and adversarial training. The third axis, semantic scope, ranges from single explicit concepts to multi-concept and concept-combination erasure. The mechanism the taxonomy is supposed to carry is comparison: by locating any method in this space, a reader can see which architectural site, learning strategy, and target complexity it uses, and the two qualitative tables translate that location into expected effectiveness and utility.","core_discovery":"On its own terms, the paper's central contribution is a three-dimensional classification framework for concept erasure techniques, with every reviewed method located by intervention level (text encoder, cross-attention, UNet), optimization strategy (loss-based, closed-form, plug-in adapters, adversarial training), and semantic scope (explicit, multi-concept, concept combination). It further claims that this organization reveals systematic trade-offs—for instance, that adversarial training gives the highest erasure effectiveness at the cost of low utility, while plug-in adapters preserve utility at moderate effectiveness—and that these trade-offs are consistent across intervention levels. The survey also asserts that current evaluation practice, centered on Erasure Success Rate, FID, and CLIP Score plus a growing set of benchmarks, is not yet standardized enough to support fully rigorous comparison, especially for robustness and practical deployment.","pith_inferences":["If the taxonomy is taken as a generative map, one testable extension is to use it to predict the performance of an unseen method from its coordinates, which the survey itself does not explicitly propose.","The paper rates effectiveness and utility without a shared evaluation protocol; a natural next step would be a standardized benchmark that computes ESR, FID, and CLIP Score under identical prompts and seeds for one representative method per cell of the taxonomy.","The framework's intervention-level axis is tied to the UNet/cross-attention architecture of Stable Diffusion; extending the survey to flow-matching models such as Flux, which the paper notes lack cross-attention, may require a fourth axis or a redefinition of intervention sites."],"forward_implications":["Given any candidate erasure method, a practitioner can locate it in the three axes and infer its likely trade-off profile, because the paper claims the taxonomy 'allows for clearer methodological comparisons.'","Methods that intervene more deeply (UNet-level) are predicted to achieve stronger, more specific suppression but with lower utility and higher computational cost, while shallow interventions trade robustness for efficiency and transferability.","Adversarial training consistently shows the highest effectiveness ratings but the lowest utility ratings across all intervention levels, so robust erasure should be paired with utility-preserving components.","No single method combines top erasure strength with top utility; the paper therefore argues that hybrid multi-layer, multi-objective pipelines are the promising next step.","The evaluation landscape is fragmented enough that the paper's recommended benchmarks and metrics (ESR, FID, CLIP Score, plus integrated datasets like UnlearnCanvas and HUB) will need further standardization before cross-method comparisons become truly reliable."],"supporting_citations":[{"why":"The prior survey on concept erasure that this work extends; it supplies the baseline categorization the three-dimensional taxonomy refines.","marker":"[13]"},{"why":"ESD introduces the widely used loss-based erasure objective that many methods in the taxonomy build on, anchoring the loss-based optimization category.","marker":"[14]"},{"why":"UCE is the pioneering closed-form projection method, anchoring the closed-form category through its analytical key/value updates.","marker":"[49]"},{"why":"The I2P dataset provides standard unsafe prompts used to benchmark NSFW erasure effectiveness across the field.","marker":"[91]"},{"why":"CLIP is the common text encoder whose prompt embeddings are the target of text-encoder-level interventions.","marker":"[98]"},{"why":"Stable Diffusion's latent diffusion architecture defines the UNet and cross-attention sites that structure the intervention-level axis.","marker":"[99]"}],"fun_headline_variants":["Three axes expose the cost of erasing concepts","A three-way map of every concept suppression trick","Survey reveals trade-offs in concept erasure methods","Where, how, and what: a taxonomy for safe T2I"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparative conclusions rest on the paper's own qualitative Effectiveness and Utility ratings in Tables I and II, which are not backed by a common evaluation protocol or quantitative results.","fun_headline_variants_meta":{"raw":{"variants":["Three axes expose the cost of erasing concepts","A three-way map of every concept suppression trick","Survey reveals trade-offs in concept erasure methods","Where, how, and what: a taxonomy for safe T2I"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000693,"raw_usage":{"total_tokens":3141,"prompt_tokens":958,"completion_tokens":2183,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":2131}},"tokens_in":574,"tokens_out":2183,"duration_ms":18647,"temperature":1.0,"reasoning_tokens":2131,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:15:53.724800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A standardized evaluation that runs one representative method per taxonomy cell on the same prompts, metrics, and random seeds would falsify the survey's trade-off claims if the measured Effectiveness and Utility rankings contradict the ratings in Tables I and II (for example, if a text-encoder-level loss-based method outperformed every UNet-level method in both effectiveness and utility).","supporting_citations":[{"cited_title":"Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models,","cited_arxiv_id":null,"evidence_quote":"The I2P dataset provides standard unsafe prompts used to benchmark NSFW erasure effectiveness across the field."},{"cited_title":"Learning transferable visual models from natural language supervi- sion,","cited_arxiv_id":null,"evidence_quote":"CLIP is the common text encoder whose prompt embeddings are the target of text-encoder-level interventions."},{"cited_title":"High-resolution image synthesis with latent diffusion models,","cited_arxiv_id":null,"evidence_quote":"Stable Diffusion's latent diffusion architecture defines the UNet and cross-attention sites that structure the intervention-level axis."}],"review_version":1}