{"id":"fc1c17f3-1068-4943-99cb-d13a04a2c15a","arxiv_id":"2511.19316","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Existing dataset watermarks for diffusion fine-tuning transfer well across models and tasks but remain vulnerable to a proposed restoration-based removal attack (DeAttack), whose claimed full removal is not fully demonstrated.","lead":"This paper benchmarks four dataset-watermarking methods for tracing images used to fine-tune diffusion models, measuring universality, transmissibility, and robustness. It then proposes a restoration-based removal attack called DeAttack and argues that current watermarks survive common distortions but not targeted removal.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DeAttack removal claim is unsupported: Table 5's caption promises nine 'gray-shaded' DeAttack columns, but the printed table contains only six baseline restorers; the visible columns show DiffusionShield at 98–100% accuracy, contradicting the abstract's 'completely remove' assertion.","rationale":"The reader flagged the undisclosed watermark embedding-strength calibration as the weakest assumption, which is a valid fairness concern for the benchmark rankings. My stress test identifies an even more direct problem: the paper's central removal claim is not backed by the presented evidence. Table 5 is explicitly described as including DeAttack results ('last nine gray-shaded columns'), but those columns are missing from the manuscript, and the visible baseline results show DiffusionShield at ~100% detection under all six restorers. This is an internal inconsistency, not a matter of interpretation: either the table was truncated/corrupted in the manuscript, or the results were never obtained. Either way, the abstract's assertion that existing methods 'fall short under real-world threat scenarios' and that DeAttack 'can completely remove dataset watermarks' is not supported by the data shown. The benchmark contribution (Tables 2–4) may still be valuable, and the calibration concern should also be addressed, but the removal claim requires at minimum the full DeAttack table and a demonstration that at least the most robust method (DiffusionShield) drops to chance without degrading generation quality. The current evidence actually points in the opposite direction. I therefore recommend the same overall verdict as the reader, CONDITIONAL, but with the condition sharpened: the paper should be accepted (if at all) only after the complete Table 5 is provided and the removal claim is verified; otherwise the central contribution is unverified and the abstract overstates the findings.","tokens_in":14612,"tokens_out":2748,"duration_ms":29321,"concrete_test":"Obtain the full Table 5 from the authors or the released code, including the promised nine gray-shaded DeAttack columns. Verify whether these columns exist and report detection accuracy, FID, and CLIP-T for each watermarking method after DeAttack. The claim 'completely remove' requires that Acc drop to near chance (e.g., ~50% for SIREN-like baselines and DiffusionShield) while FID and CLIP-T remain close to the clean fine-tuning baseline (e.g., within the range of the Clean rows in Table 2). If DiffusionShield accuracy remains above 90% in any DeAttack column, the central claim is falsified; if the columns are absent, the claim is unverified and should be withdrawn unless the evidence is supplied.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline contribution is DeAttack, a watermark removal method claimed to \"completely remove dataset watermarks without affecting fine-tuning.\" This claim rests on the experimental evidence in Table 5, but that table is internally inconsistent and does not report DeAttack results. The caption states \"The last nine gray-shaded columns correspond to our methods,\" yet the rendered table contains only six columns (Bmshj2018, Cheng et al., Diffusion, SwinIR Denoise, SwinIR JPEG AR, IRNeXt Deblur), none of which are identified as DeAttack. Section 5 then states that \"our IRNeXt-based DeAttack achieves stronger watermark removal with minimal perceptual distortion, as validated in Table 5,\" but no DeAttack column exists to validate this. Moreover, the six columns that are shown contradict the \"completely remove\" claim: DiffusionShield detection accuracy remains 99.99%, 99.78%, 100.00%, 98.39%, 100.00%, and 100.00% across the six baseline restorers, and WatermarkDM remains at 48.44–62.50%. Thus, on the only evidence presented, at least one method is not removed by any of the listed degradation-restoration pipelines. The abstract's absolute removal claim is therefore not established by the manuscript; it is either missing from the tables or refuted by the data shown. Because the removal approach is the paper's stated novel contribution, this is the most load-bearing weakness: without DeAttack-specific results, the central claim is unsupported, and the robustness conclusions in the abstract are overreaching.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a unified threat model and evaluation framework for dataset watermarking in customized diffusion models, benchmarking four watermarking methods (DIAGNOSIS, DiffusionShield, SIREN, WatermarkDM) across three datasets, four fine-tuning approaches, two text-encoder settings, several watermarked-subset ratios, and three common distortions. It reports that existing methods are largely universal and transmissible, with DiffusionShield the most robust and SIREN near chance. Based on this evaluation, the paper introduces DeAttack, a degradation-restoration removal framework claimed to 'completely remove dataset watermarks without affecting fine-tuning.' The benchmark portion is substantial, but the removal claim is not supported by the evidence actually presented in Table 5, which contains no DeAttack-specific columns and shows DiffusionShield retaining near-perfect detection accuracy under all six listed restoration baselines.","tokens_in":15052,"tokens_out":3191,"duration_ms":34938,"significance":"If the benchmark is taken at face value, it provides a useful comparison resource: it covers four open-source watermarking methods, multiple fine-tuning paradigms, three datasets, and a range of practical distortions, with FID, CLIP-T, and detection accuracy reported. Using each method's own watermark decoder is appropriate for a traceability test, so the 'circularity' concern is not the main issue. However, the paper's stated novel contribution—DeAttack—is not substantiated: the table that is supposed to validate it appears to report only off-the-shelf restoration baselines, and the visible numbers directly contradict the abstract's 'completely remove' claim. The calibration of watermark embedding strength is also undisclosed, which threatens the fairness of cross-method comparisons. The benchmark has real value, but the removal contribution needs substantial rework or repositioning.","major_comments":[{"comment":"The caption of Table 5 states that 'the last nine gray-shaded columns correspond to our methods,' but the rendered table contains only six columns—Bmshj2018, Cheng et al., Diffusion, SwinIR (Denoise), SwinIR (JPEG AR), and IRNeXt (Deblur)—none of which is identified as DeAttack. Section 5 then claims that 'our IRNeXt-based DeAttack achieves stronger watermark removal ... as validated in Table 5,' but no DeAttack column exists to validate this. Moreover, the six columns that are shown directly contradict the abstract's 'completely remove' assertion: DiffusionShield detection accuracy remains 98.39–100.00%, and WatermarkDM remains 48.44–62.50%. On the evidence presented, at least one method is not removed by any of the listed degradation-restoration pipelines. The authors must either supply the actual DeAttack results (including the proposed degradation-restoration composition and ablation","section":"Table 5 and Section 5"},{"comment":"The statement 'we calibrate the watermark embedding strength across all methods' is load-bearing for every ranking in Tables 2–5, but no calibration procedure, objective, target utility range, or per-method embedding strengths are reported. Since detection accuracy is monotone in embedding strength for most watermarking schemes, an arbitrary or method-specific calibration could determine the observed ranking (e.g., DiffusionShield near 100% vs. SIREN near 50%). Without this information, the fairness of the benchmark and the robustness conclusions do not follow. Please disclose the calibration protocol, the actual strength values, and the resulting FID/CLIP-T ranges.","section":"Section 4.2, Universality Evaluation"},{"comment":"DeAttack is described as a 'unified framework' and an optimization in Eq. (2), but the implementation described in the text is simply the composition of pretrained restoration models: IRNeXt trained on DIV2K/Flickr2K/WED and two pretrained SwinIR models. There is no description of how Eq. (2) is solved, what the learnable degradation prior P_θ in Eq. (13) is, or how the proposed autoencoder in Figure 4 is trained or adapted. If DeAttack is just an application of existing restoration networks, then the claim of a new removal method is overstated, and the benchmark in Table 5 is not an evaluation of DeAttack at all. Please clarify the novelty and provide a concrete algorithmic description.","section":"Section 5, Eqs. (2)–(14)"},{"comment":"All detection accuracy numbers are obtained with each method's original detector and threshold. After degradation and restoration, the distribution of generated images changes substantially, and a fixed threshold may no longer be statistically calibrated. The paper should at least report ROC curves or detection scores at a fixed false-positive rate, rather than only accuracy at a threshold that may not correspond to comparable operating points across methods. This is secondary to the missing DeAttack results, but it affects the robustness comparison as well.","section":"Tables 4 and 5, detection metric"}],"minor_comments":[{"comment":"The text says 'Compared to the results obtained when fine-tuning with fully watermarked images in Table 3,' but Table 3 reports mixed-ratio results; the fully watermarked results are in Table 2. Please correct the cross-reference.","section":"Section 4.3"},{"comment":"The operator E(·) in Eq. (2) is referred to as 'a watermark extractor or spectral energy operator,' but no concrete definition, implementation, or choice of λ is given. This makes the proposed optimization untestable as stated.","section":"Section 5, Eq. (2)"},{"comment":"The caption says the figure indicates 'the optimal FID score and CLIP-T similarity for each fine-tuning approach,' but it is unclear what 'optimal' means here and how the figure should be read. A legend or explanation of the displayed samples would help.","section":"Figure 3"},{"comment":"The table header labels the columns 'Results of different DeAttack methods,' but the columns are named after well-known baseline restorers. Even if the authors intended these as DeAttack variants, the naming is misleading and should be clarified.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The benchmark portion of this manuscript is potentially publishable, but the central removal contribution is currently unsupported and appears contradicted by the data in Table 5. The authors should either provide the missing DeAttack results or restructure the paper as a benchmark-only study. The calibration disclosure in Section 4.2 is mandatory before the cross-method rankings can be trusted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark part of this paper is the real contribution. A unified three-axis evaluation of DIAGNOSIS, DiffusionShield, SIREN, and WatermarkDM across four fine-tuning methods, three datasets, several protection ratios, and common distortions is genuinely new and useful. The transmissibility results at 20–80% protection ratios are the kind of empirical data this field needs. I don't see an obvious problem with using each method's own detector for a traceability evaluation, and the threat model is sensible.\n\nThe soft spot is exactly where the abstract makes its strongest claim. Table 5's caption promises nine gray-shaded DeAttack columns, but the actual table shows only six baseline restorers—Bmshj2018, Cheng, Diffusion, SwinIR denoise, SwinIR JPEG AR, IRNeXt deblur. There is no DeAttack column anywhere. Section 5 says the IRNeXt-based DeAttack is validated in Table 5, but that validation is not visible. Worse, the six columns that do appear show DiffusionShield detection accuracy at 98–100% across all restorers, which directly contradicts the abstract's \"completely remove\" statement. This is a load-bearing gap, not a quibble: the removal method is the paper's stated novel contribution, and the evidence for it is either missing or points the other way.\n\nA second issue, less glaring but still important: the fairness calibration in Section 4.2 is opaque. The authors say they calibrated watermark embedding strength across methods to make FID/CLIP-T comparable, but give no procedure or values. Since detection accuracy depends strongly on embedding strength, the cross-method rankings in Tables 2–4 are only as trustworthy as that undisclosed calibration. This needs to be specified in a revision.\n\nWho is this for? Anyone working on dataset watermarking for customized diffusion models or generative-model copyright protection. The benchmark deserves to exist and be cited. But the paper in its current form overreaches: the removal claim should either be supported with actual DeAttack results across datasets and fine-tuning methods, or removed from the abstract. I'd send it to peer review because the benchmark is valuable enough to warrant referee time, but I'd expect major revision and would want the missing data and calibration details before believing the headline.","headline":"The benchmark work is real and useful, but the DeAttack removal claim is not backed by the supplied table—the promised DeAttack columns are missing and the baselines shown leave DiffusionShield at ~100% detection.","tokens_in":15468,"tokens_out":1760,"would_cite":true,"duration_ms":20479,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that existing dataset watermarking methods for diffusion-model fine-tuning are universal and transmissible but can be completely removed by a proposed degrade-and-restore attack, DeAttack, which erases watermarks without h","keywords":["dataset watermarking","diffusion model","fine-tuning traceability","watermark removal","benchmark","threat model","robustness","customized generation"],"falsifier":"Re-run the benchmark with DeAttack applied to the watermarked datasets, fine-tune with the same protocols, and measure detection accuracy: if any method (for example, DiffusionShield) keeps accuracy near 100% after removal, or if fine-tuning utility degrades sharply in FID or CLIP-T, the paper's central removal claim fails. Also, if the embedding-strength calibration values are disclosed and re-running with different calibration changes the method rankings, the fairness premise fails.","tokens_in":14547,"feed_emoji":"🖼️","tokens_out":3624,"duration_ms":36752,"temperature":0.7,"pith_summary":"The paper argues that current dataset watermarking methods for tracing unauthorized fine-tuning of diffusion models are not ready for real-world adversaries. It builds a unified benchmark testing three properties—universality across fine-tuning methods and tasks, transmissibility when only part of the dataset is watermarked, and robustness to post-processing—and finds the methods pass the first two but fail the third. The paper then introduces DeAttack, a watermark removal method that degrades watermarked images and restores them, claiming it completely eliminates dataset watermarks while leaving fine-tuning performance intact. If true, this means existing protection can be bypassed by a determined user, and watermark designers must treat degradation-restoration as a core threat.","feed_headline":"New attack erases dataset watermarks from diffusion models","feed_subtitle":"Benchmark of four methods shows universal fine-tuning traceability is undone by a degradation-restoration attack.","key_machinery":"The key mechanism is the regeneration attack formulated as x' = R(D(x)), where D degrades the watermarked image to disrupt the embedded watermark and R restores perceptual quality to make the result usable for fine-tuning. DeAttack instantiates this with an autoencoder-based pipeline combining pixel-space and latent-space degradations with restoration networks such as IRNeXt and SwinIR, trained on standard image restoration data. This orchestrated degrade-then-restore loop is what breaks watermarks that survive ordinary noise, blur, or JPEG.","core_discovery":"The central claim is that dataset watermarking methods for diffusion models, while universal and transmissible, cannot withstand a targeted degradation-restoration removal attack. Under the benchmark, DiffusionShield achieves near-perfect detection across all fine-tuning settings, while SIREN performs near random chance; the other methods sit in between and often collapse when the text encoder is trained. The paper's DeAttack applies degradation (Gaussian noise, blur, JPEG-like compression, latent-space noise) followed by high-quality restoration with pretrained networks, reducing watermark detection accuracy to chance without measurable harm to generation quality. The conclusion is that rob","pith_inferences":["One extension, not pursued here, is to test watermarking schemes that embed signals in model weights or prompts rather than training images; these may resist degradation-restoration attacks that target pixel-level watermarks.","The benchmark's structure suggests that robustness should be measured against an adaptive adversary who knows the watermarking method, since static robustness tests underestimate real-world risk.","DeAttack could be repurposed as a data augmentation for training more robust watermarks, turning the removal attack into a defense that forces watermarks to survive degradation and restoration.","The paper's finding implies that any practical deployment of dataset watermarking should assume an adversary will apply some form of input purification, and should therefore target semantic or structural features that survive such purification."],"forward_implications":["If DeAttack works as claimed, dataset watermarking cannot by itself guarantee traceability against a determined adversary who applies degradation and restoration before fine-tuning.","Robustness evaluations of watermarking methods should include adversarial removal attacks, not only common image distortions, or they will overstate real-world protection.","Among the tested methods, DiffusionShield is the strongest candidate for practical use, while SIREN provides essentially no traceability in these settings.","Training the text encoder during fine-tuning weakens several watermarking methods, suggesting text-encoder updates are a natural attack surface.","The proposed threat model and three-dimensional benchmark provide a shared protocol for comparing future dataset watermarking methods."],"fun_headline_variants":["Watermark removal attack breaks diffusion model traceability","Benchmark exposes diffusion watermark weakness, offers removal attack","Degradation-restoration erases diffusion watermarks, study shows","Diffusion watermarking fails against new removal method","Novel attack removes dataset watermarks from fine-tuned diffusion models"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire comparison rests on an undisclosed calibration of watermark embedding strength across methods, and on treating each method's own detector and threshold as comparable; if that calibration is arbitrary, the detection-accuracy rankings and the claimed robustness conclusions do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Watermark removal attack breaks diffusion model traceability","Benchmark exposes diffusion watermark weakness, offers removal attack","Degradation-restoration erases diffusion watermarks, study shows","Diffusion watermarking fails against new removal method","Novel attack removes dataset watermarks from fine-tuned diffusion models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000454,"raw_usage":{"total_tokens":2075,"prompt_tokens":653,"completion_tokens":1422,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":1343}},"tokens_in":397,"tokens_out":1422,"duration_ms":10588,"temperature":1.0,"reasoning_tokens":1343,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T20:30:31.303807+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the benchmark with DeAttack applied to the watermarked datasets, fine-tune with the same protocols, and measure detection accuracy: if any method (for example, DiffusionShield) keeps accuracy near 100% after removal, or if fine-tuning utility degrades sharply in FID or CLIP-T, the paper's central removal claim fails. Also, if the embedding-strength calibration values are disclosed and re-running with different calibration changes the method rankings, the fairness premise fails.","supporting_citations":[],"review_version":1}