{"id":"0efe7f3d-47f6-4f5b-b1d3-a0eb05602eed","arxiv_id":"2504.18032","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"PRSS replaces the unconditional anchor in classifier-free guidance with the original prompt and searches for semantically similar alternative prompts, improving the privacy-utility trade-off against memorization baselines.","lead":"This paper introduces PRSS, a way to reduce how often Stable Diffusion reproduces memorized training images by changing how the model's guidance signal is computed at inference time. It combines two techniques: anchoring generation to the original prompt instead of the empty prompt, and using an LLM to find safer paraphrases of the prompt.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aggregate trade-off curves without error bars or detector-outcome stratification do not yet establish the claimed consistent Pareto improvement.","rationale":"I read the paper's central claim as an empirical dominance claim: PRSS's inference-time CFG modification is claimed to be a strict improvement over prompt-engineering baselines. I looked for the weakest link in that chain. The method description is clear, and the mechanism of using the original prompt as an anchor plus an LLM-selected paraphrase is plausible; the ablations and supplementary metrics support the intended direction. I do not see an internal inconsistency in Eq. (12). The main risk is that the evidence for dominance is aggregate and unquantified. The reader flagged detector accuracy as the weakest assumption and also noted missing error bars and code; I partially agree. I would sharpen the concern as statistical reliability and detector-outcome stratification, because detector accuracy alone is a shared limitation with the baselines, whereas missing uncertainty quantification and per-outcome analysis directly affect whether 'consistently improves' is supported. A multi-seed rerun with confidence intervals and oracle-triggered stratification would settle the concern. This does not change the conditional verdict: the method may well work, but the current manuscript does not yet establish robustness of the headline trade-off claim.","tokens_in":17657,"tokens_out":10973,"duration_ms":123765,"concrete_test":"Re-run the Fig. 6 protocol with at least 10 random seeds per (method, detection signal, lambda) and report mean and 95% confidence intervals for CLIP and SSCD; then stratify the same data by detector outcome (true positives, false negatives, false positives) or, better, replace the detector by an oracle trigger using ground-truth memorization labels. If PRSS is not Pareto-dominant within the confidence intervals on the oracle-triggered subset, or if its apparent dominance vanishes on true positives, the consistent-improvement claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"PRSS's central claim is empirical: Eq. (12) gives higher CLIP at matched SSCD than the prompt-engineering baselines. The supporting evidence is Figs. 6, 7, and 9-11, which show aggregate curves but no error bars, no number of seeds, no per-prompt scatter, and no numerical table. Diffusion sampling is stochastic; without repeated-seed intervals, a 'consistent' Pareto dominance could be within seed noise, especially for local memorization where the paper itself reports a smaller margin. The curves are also aggregates over the detection gate: at high lambda many memorized prompts are false negatives and both methods revert to standard Stable Diffusion, while at low lambda many safe prompts are false positives and are modified unnecessarily. If PRSS's advantage comes from a different distribution of gated prompts rather than from a better per-prompt trade-off, the headline claim is not established. The conclusion acknowledges detector dependence, but the paper never quantifies it or separates detector error from mitigation quality. This is a missing-evidence problem rather than a logical contradiction, but it is the load-bearing support for the SOTA claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PRSS, an inference-time memorization mitigation method for text-to-image diffusion models. PRSS refines classifier-free guidance: when a first-step magnitude signal indicates memorization risk, the usual null anchor of CFG is replaced by the original user-prompt embedding (prompt re-anchoring, PR), and the conditional embedding is replaced by an LLM-searched semantically similar alternative (semantic prompt search, SS). The method is designed to plug into existing detection signals from Wen et al. [36] and Chen et al. [8]. The authors evaluate on Stable Diffusion v1-4 using the 500-prompt dataset of Webster [34], reporting privacy-utility trade-off curves (CLIP score vs. SSCD average, SSCD 95th percentile, LS, and percent memorized) for global and local memorization, with comparisons against the two prompt-engineering baselines and ablations of PR and SS. They claim consistent improvement over the baselines and a new state of the art.","tokens_in":17855,"tokens_out":4434,"duration_ms":45207,"significance":"If the empirical claim holds, PRSS is a practically attractive contribution: it requires no retraining or training-data search, alters only the guidance equation at inference, and leverages a cheap LLM-based prompt search. The conceptual move of using the memorized prompt itself as a negative anchor is intuitive and distinct from prior prompt-engineering approaches. However, the paper's central claim is empirical, and the reported evidence is currently insufficient to establish a consistent, statistically reliable Pareto improvement. The strength of the paper is its simple and seemingly effective recipe; the weakness is the absence of repeated-seed uncertainty quantification, stratification by detector outcome, and transparency about baseline reimplementation. The authors also explicitly acknowledge sensitivity to the detection mechanism, which is a known limitation but one that the experiments do not quantify. The result is plausible but not yet convincingly supported as stated.","major_comments":[{"comment":"The central claim of a consistent privacy-utility improvement is supported only by aggregate trade-off curves without error bars, without the number of seeds used, and without significance tests. Diffusion sampling is stochastic, and the curves in Fig. 6 show a smaller margin for local memorization, so the observed dominance could in principle be within seed noise. The authors should report repeated-seed intervals (e.g., mean and variance over 3-5 seeds) or per-prompt scatter plots, and where possible a paired significance test (e.g., Wilcoxon signed-rank) at matched privacy levels.","section":"Sec. 4.2, Figs. 6-7 and 9-11"},{"comment":"The trade-off curves aggregate over the detection gate: at a given threshold lambda, some memorized prompts are false negatives (reverting to standard Stable Diffusion) and some safe prompts are false positives (modified unnecessarily). Because of this, a shift in the curve could reflect a different distribution of gated prompts rather than a better per-prompt mitigation trade-off. The conclusion acknowledges the dependence on detector accuracy, but the paper never quantifies detector false-positive and false-negative rates at the thresholds used, nor reports results stratified by detection outcome. The authors should report detector accuracy on the evaluated prompt set and present trade-off curves or tables for true-positive, false-negative, false-positive, and true-negative subsets separately.","section":"Sec. 4.2 and Sec. 6"},{"comment":"The baselines from [36] and especially from [8] are self-reimplemented, and the paper does not state whether official code or checkpoints were used or how hyperparameters were matched. Since the strongest comparison is against the authors' own prior work [8], the reader needs more detail on the reimplementation (e.g., exact prompt-engineering loss, optimization steps, and stopping criterion) and on the lambda values used to produce each curve point. Without this, the fairness of the comparison and the reproducibility of the trade-off curves are hard to verify.","section":"Sec. 4.1 and Figs. 6-11"}],"minor_comments":[{"comment":"The definition of LS is notationally awkward: the indicator 1_{SSCD>0.5} should be a scalar factor multiplying the norm, not a dot product, and the subscript is typeset without an underscore.","section":"Eq. (6)"},{"comment":"The denominator in the definition of m'_t is garbled in the text ('N~ 1 N Nÿ i=1 mi'); please rewrite it as a clear normalized sum, e.g., ||...||_2 / (1/N * sum_i m_i).","section":"Eq. (8)"},{"comment":"The adaptive guidance strength s_1 is said to depend on 'the real-time magnitude at timestep t', but it is not specified which magnitude (e.g., between e_p and e_phi, or between e^{ss}_p and e_p) is used at later timesteps. Please define m_t in this context and clarify whether it is recomputed at every denoising step.","section":"Sec. 11.1, Eqs. (15)-(16)"},{"comment":"The text says the search generates 'up to n_s = 25 semantically similar prompts' in Sec. 4.1, but the implementation details say 'n=1 generates a single prompt per call'; please reconcile the notation and explain how the 25 alternatives are collected.","section":"Sec. 11.2"},{"comment":"The phrase 'astonishingly large improvements' is informal for a journal report; please replace it with a quantitative statement, e.g., the observed reduction in average SSCD at matched CLIP score.","section":"Sec. 4.2"},{"comment":"The paper states that 'over 300' of the 500 prompts are memorized; please give the exact counts for the global and local memorization subsets, as these are the denominator for the reported percentages in Fig. 11.","section":"Sec. 4.1"}],"recommendation":"major_revision","confidential_remarks":"The strongest baseline [8] and the detection signals used throughout are from the authors' own prior work, which makes independent replication particularly important; I would encourage the editor to require the authors to release code or at least detailed per-prompt results. The missing uncertainty quantification is the main technical gap, but it is fixable within the scope of the manuscript, so I do not see a basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version. PRSS is a clean, cheap inference-time idea: when memorization is detected, re-anchor classifier-free guidance on the original prompt instead of the empty prompt, and use an LLM to substitute a semantically similar prompt. That combination is new relative to the authors' own prior prompt-engineering pipeline, and the reported trade-off curves consistently sit in its favor. But the 'new state-of-the-art' claim is not actually established, because the evidence consists of aggregate curves with no error bars, no seed counts, and no breakdown by detector outcome.\n\nWhat the paper does well: the method is simple and requires no retraining; it plugs into any detection signal (m or m'); the ablation shows both components earn their keep; and the evaluation uses multiple privacy metrics (SSCD average, 95th percentile, LS, percent memorized) plus CLIP for utility, so the result is not circular. The geometric analysis is just intuition, but it's reasonable intuition and the authors don't oversell it as theory.\n\nThe soft spots are real. The stress-test note is on target: the main claim is empirical and the evidence is under-specified. Diffusion sampling is stochastic; without repeated-seed intervals, a 'consistent' Pareto improvement could be within seed noise, especially for local memorization where the paper itself reports a smaller margin. The method also gates on a memorization detector; the conclusion explicitly acknowledges the dependence, but the paper never quantifies false negatives/positives or shows the trade-off curves conditioned on detector outcomes. If PRSS's advantage comes from triggering on a different set of prompts than the baseline, the Pareto claim is not yet supported. Baselines are self-reimplemented and no code or data are released, so an independent check is hard. These are fixable problems, not contradictions.\n\nWho should read it: anyone working on memorization mitigation in text-to-image models. It deserves a serious referee; the idea is worth engaging, but the empirical section needs strengthening. My recommendation: send it to review, and in the review ask for error bars, seed counts, detector-outcome-stratified curves, and code release.","headline":"PRSS is a plausible inference-time tweak to CFG, but the SOTA claim outruns the evidence: no error bars, no seeds, no detector-outcome stratification.","tokens_in":18394,"tokens_out":2994,"would_cite":true,"duration_ms":29072,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Changing which conditioning vector anchors classifier-free guidance can reduce training-image copying in diffusion models while preserving prompt alignment.","keywords":["memorization","diffusion models","classifier-free guidance","privacy-utility trade-off","prompt re-anchoring","semantic prompt search","text-to-image generation","SSCD similarity"],"falsifier":"Build a test set of prompts that are known to reproduce training images but whose first-step conditional-versus-unconditional magnitude stays below the threshold; if such prompts are easy to find, PRSS would produce the same memorized images as the unmodified model, showing that the method's protection is exactly as good as the detector.","tokens_in":44,"feed_emoji":"🖼️","tokens_out":9010,"duration_ms":145676,"temperature":0.7,"pith_summary":"Text-to-image diffusion models can reproduce images from their training sets, creating privacy and copyright risk, but the usual fix of rewriting the user prompt to suppress recall also degrades how well the output matches the prompt. This paper claims the privacy-utility trade-off can be improved by changing the geometry of classifier-free guidance instead of only editing the text. Its PRSS method keeps the original prompt as the guidance anchor and pulls the image away from that anchor toward a paraphrased, semantically similar prompt found by a large language model. In experiments on known memorized prompts, the method reports lower copy-detection similarity at the same text-alignment level as prompt-engineering baselines, for both global and local memorization. If the claim holds, model operators gain an inference-only lever that reduces copying without retraining or training-data access.","feed_headline":"Re-anchored prompts cut memorization and keep alignment","feed_subtitle":"PRSS re-anchors on the original prompt and swaps in a paraphrased one, beating prompt editing on both axes.","key_machinery":"The load-bearing object is a modified classifier-free guidance equation, Eq. (12). In standard CFG the noise prediction is pulled from a null-text prediction toward the prompt-conditioned prediction; PRSS, on detecting memorization, instead anchors at the original prompt's prediction and points toward a semantic alternative's prediction. The detector is the first-step magnitude $m_{T-1} = \\|\\epsilon_\\theta(x_{T-1},e_p) - \\epsilon_\\theta(x_{T-1},e_{\\phi})\\|_2$ inherited from prior work, optionally masked to target local memorization. The semantic alternative is found by asking a large language model for paraphrases and picking the first whose magnitude falls below the threshold. The mechanism's work is to make the guidance direction itself privacy-preserving, so the prompt does not have to be distorted to suppress copying.","core_discovery":"The paper's central claim is that both legs of the classifier-free guidance update are suboptimal for privacy: the text-conditional leg over-edits the prompt, and the unconditional null anchor does not point away from memorized content. PRSS therefore replaces the null anchor $e_{\\phi}$ with the original prompt embedding $e_p$ and replaces the engineered prompt $e^*$ with a language-model-searched semantic alternative $e^{ss}_p$, producing the modified update in Eq. (12). When the first-step magnitude signal exceeds a threshold, PRSS generates along the direction $\\epsilon_\\theta(x_t,e_p) + s(\\epsilon_\\theta(x_t,e^{ss}_p) - \\epsilon_\\theta(x_t,e_p))$; otherwise it keeps standard CFG. The reported experiments show this re-anchored guidance yields higher CLIP alignment at matched SSCD similarity than prompt engineering, and pairing it with the stronger masked detection signal gives the best results. The authors state that this consistently improves the privacy-utility trade-off, establishing a new state of the art.","pith_inferences":["Beyond the paper: the same re-anchoring move could be applied to any undesired conditioning direction, such as negative prompts, copyrighted styles, or protected attributes, making the choice of CFG anchor a general design axis rather than a memorization-only fix.","Beyond the paper: because PRSS inherits its trigger from the detection signal, any improvement in first-step memorization detection accuracy should translate directly into better mitigation, making detection and mitigation complementary rather than competing.","Beyond the paper: with the reported search cost of about 0.9 seconds and $0.02 per prompt, alternatives could be cached or generated offline for frequent prompts, making the per-generation overhead negligible in deployment."],"forward_implications":["At a matched level of copy-detection similarity, PRSS reports higher CLIP text-alignment than prompt-engineering baselines on both global and local memorization prompts.","Combining PRSS with the stronger masked detection signal yields the best privacy-utility trade-off, and even the weaker detection signal lets PRSS match the baseline that uses the stronger signal.","Ablations show prompt re-anchoring alone raises privacy at the cost of utility, semantic search alone raises utility with limited privacy, and together they dominate the baseline trade-off curve.","Prompt re-anchoring keeps diverting generation across the whole denoising trajectory, preventing memorization from reappearing in later steps.","The method is inference-only: it requires no retraining, fine-tuning, or search over the training set, only a modified CFG noise-prediction equation."],"supporting_citations":[{"why":"Supplies the first-step magnitude detection signal $m_{T-1}$ and the prompt-engineering mitigation baseline that PRSS modifies and outperforms.","marker":"[36]"},{"why":"Introduces the masked magnitude $m'_t$ and local memorization masks; PRSS uses this as the stronger detection signal and as the current baseline for comparison.","marker":"[8]"},{"why":"Defines the classifier-free guidance update that PRSS re-anchors during inference.","marker":"[12]"},{"why":"Provides the prompt dataset, with over 300 memorized prompts, on which the experiments and global/local splits are based.","marker":"[34]"},{"why":"Defines the SSCD embedding and copy-detection similarity used as the privacy metric for measured memorization.","marker":"[25]"},{"why":"Establishes the object-level definition of memorization and the SSCD-based evaluation convention that the paper's metrics follow.","marker":"[31]"}],"fun_headline_variants":["PRSS: better privacy-utility, state-of-the-art","Re-anchored guidance beats prompt editing","Semantic search keeps output aligned while cutting memorization","New diffusion technique improves privacy-utility trade-off","PRSS: re-anchor to forget, search to stay on track"],"cache_read_input_tokens":20608,"weakest_assumption_plain":"The whole method switches on a first-step magnitude signal: if that signal fails to flag a memorized prompt, PRSS does nothing and generation follows ordinary Stable Diffusion.","fun_headline_variants_meta":{"raw":{"variants":["PRSS: better privacy-utility, state-of-the-art","Re-anchored guidance beats prompt editing","Semantic search keeps output aligned while cutting memorization","New diffusion technique improves privacy-utility trade-off","PRSS: re-anchor to forget, search to stay on track"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1748,"prompt_tokens":910,"completion_tokens":838,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":526,"completion_tokens_details":{"reasoning_tokens":759}},"tokens_in":526,"tokens_out":838,"duration_ms":7642,"temperature":1.0,"reasoning_tokens":759,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:26:00.917335+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build a test set of prompts that are known to reproduce training images but whose first-step conditional-versus-unconditional magnitude stays below the threshold; if such prompts are easy to find, PRSS would produce the same memorized images as the unmodified model, showing that the method's protection is exactly as good as the detector.","supporting_citations":[{"cited_title":"De- tecting, explaining, and mitigating memorization in diffusion models","cited_arxiv_id":null,"evidence_quote":"Supplies the first-step magnitude detection signal $m_{T-1}$ and the prompt-engineering mitigation baseline that PRSS modifies and outperforms."},{"cited_title":"Exploring local memorization in diffusion models via bright ending attention","cited_arxiv_id":null,"evidence_quote":"Introduces the masked magnitude $m'_t$ and local memorization masks; PRSS uses this as the stronger detection signal and as the current baseline for comparison."},{"cited_title":"A self-supervised descriptor for image copy detection","cited_arxiv_id":null,"evidence_quote":"Defines the SSCD embedding and copy-detection similarity used as the privacy metric for measured memorization."},{"cited_title":"Diffusion art or digital forgery? investigating data replication in diffusion models","cited_arxiv_id":null,"evidence_quote":"Establishes the object-level definition of memorization and the SSCD-based evaluation convention that the paper's metrics follow."}],"review_version":1}