{"id":"e386a489-e09c-497a-aa4a-32c25d608d99","arxiv_id":"2411.18479","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematization of knowledge on watermarking for AI-generated content, unifying definitions, threat models, evaluation methods, and representative schemes across modalities.","lead":"This paper surveys how watermarking can mark text, images, audio, and video as AI-generated, and it defines formal properties such as detection accuracy, robustness, and forgery resistance. It serves as a reference for researchers building watermarks and for policymakers deciding how to regulate AI content.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The SoK's comprehensiveness claim rests on a non-reproducible, unlisted paper selection; without a systematic corpus check, the taxonomy and Table 1 may generalize from a biased sample.","rationale":"The paper is a SoK, not a new construction, so its correctness is largely a matter of whether its formalization and comparison faithfully represent the field. The reader's weakest assumption—representativeness of the unlisted selection—is exactly the load-bearing point: if the corpus is biased, the 'comprehensive overview' claim fails even if every individual definition is correct. I considered internal technical issues (e.g., the breadth of the robustness definition in 3.5 or the claimed impossibility of robustness to input-independent channels) but these are either standard or explicitly delegated to cited works, so they do not threaten the central claim as directly. The paper has strengths: definitions are mostly faithful restatements of peer-reviewed results, the bibliography is broad, and the open-problems section acknowledges remaining gaps. However, the absence of a reproducible selection protocol is a genuine methodological gap for a SoK; the proposed systematic corpus comparison would settle it. Thus I agree with the reader's conditional verdict and see no reason to move it.","tokens_in":25632,"tokens_out":6881,"duration_ms":61167,"concrete_test":"Run a reproducible literature search (e.g., DBLP or Semantic Scholar) for English papers up to 2024 with 'watermark' in title/abstract combined with 'generative', 'large language model', 'diffusion model', or 'AI-generated', filtered to top security/AI/multimedia venues and arXiv preprints; deduplicate and compile the full corpus. Compare this corpus against the paper's 100+ cited works and check whether any omitted paper introduces a property, attack class, evaluation metric, or modality not covered in Sections 3–5, or would change a qualitative rating in Table 1 (e.g., a video-generation watermark with provable robustness). If omissions are empty or do not affect the taxonomy, the comprehensiveness claim stands; otherwise Section 1.1 must be revised to delineate scope and selection criteria.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1.1 states the survey covered 'over 100 peer-reviewed papers' from selected venues and arXiv, but gives no search protocol, no inclusion/exclusion rules, and no paper list. The central claim—'comprehensive overview' and 'formalize the definitions and desired properties'—therefore depends on the unstated representativeness of this selection. If important schemes or paradigms were omitted or misweighted, the formalization in Section 3 and the qualitative ratings in Table 1 could mislead rather than organize the field. The risk is amplified because several core definitions (undetectability, robustness, unforgeability, distortion-freeness) are drawn primarily from the authors' own prior work ([28], [68], [73], [75], [30]); this can over-weight a single research program and under-represent, say, purely empirical post-hoc schemes or commercial deployments. The paper does provide real value—formal definitions largely restate peer-reviewed results, the bibliography is broad, and Section 7 openly lists open problems (e.g., robust + unforgeable public attribution)—so the issue is scope evidence, not internal soundness.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This SoK surveys watermarking for AI-generated content across text, image, audio, and video. It motivates watermarking via limitations of post-hoc detection and recent regulatory developments, then formalizes six desired properties (quality, false positive rate, robustness, unforgeability, message support, computational efficiency), describes threat models and evaluation practices, and reviews representative schemes in each modality. The paper closes with open problems such as robust unforgeable public attribution, proof-of-generation, semantic watermarks, and watermarkability of open-source models.","tokens_in":25753,"tokens_out":4853,"duration_ms":42371,"significance":"If the survey's coverage is representative, this is a valuable shared reference: it provides a coherent vocabulary for six watermark properties, connects formal definitions to concrete schemes, and discusses policy in a way that is directly useful to researchers and policymakers. Strengths include the explicit syntax for Detect/Decode/Attribute, the clear separation of robustness and unforgeability, the breadth of the bibliography, and the honest discussion of open problems and potential privacy abuses. The main risk is not internal logic but scope evidence: the selection of surveyed works is not reproducible, and several core definitions come from the authors' own prior work, so the 'comprehensive overview' claim needs stronger support before the taxonomy and Table 1 can be treated as field-level organizing principles.","major_comments":[{"comment":"The methodology states that 'over 100 peer-reviewed papers' were reviewed with qualitative selection criteria, but no search protocol, inclusion/exclusion rules, or list of surveyed papers is provided; the claim of a comprehensive overview is therefore not empirically checkable. Please add a supplementary list of all reviewed papers (e.g., an appendix table) with venue, year, modality, and the reason for inclusion/exclusion, and state the databases and search dates used.","section":"Section 1.1"},{"comment":"The sentence 'It also implies that it is impossible to learn detection, decoding, or attribution keys' is too strong as stated. Section 3.4 discusses unforgeable public attribution, where the attribution key is public by design, so such keys are trivially 'learnable' even in an undetectable scheme. Please qualify the claim (e.g., 'without access to the generation key and unless the scheme intentionally publishes the key') or define a precise notion of 'learning keys' from oracle access.","section":"Section 3.1.3"},{"comment":"In the manuscript as provided, Table 1 shows the method names and column headers but no rating markers in the cells, so the 'qualitative comparison' described in Section 6 cannot be inspected. Please ensure the table is rendered with the intended darker/lighter circles and add a note on how the ratings were assigned (e.g., from the cited papers' own claims or from independent evaluation).","section":"Table 1"}],"minor_comments":[{"comment":"The statement that 'Computational distortion-freeness is generally no weaker than statistical distortion-freeness in practice' is confusing, since computational distortion-freeness is the weaker notion; I suggest saying it is no stronger.","section":"Section 3.1.2"},{"comment":"The threat-model taxonomy is useful, but a table or diagram summarizing adversary capabilities (oracle access, key knowledge, verifier feedback) alongside examples would help readers map attacks to threat models.","section":"Section 4"},{"comment":"In the Gaussian Shading discussion, the text says the proof does not account for correlations across multiple images; please clarify that this is a gap between the proved single-response property and the multi-response quality metrics (FID, CLIP score, etc.), not a flaw in the proof itself.","section":"Section 6.2"},{"comment":"Reference [45] cites 'Exec. Order No. 1411' but the order discussed in the text is Executive Order 14110; the reference list entry should be corrected.","section":"References"},{"comment":"The table row 'Pseudorandom Codes [73]' refers to a text watermark, while Section 6.2 describes the PRC watermark [30] for images; please make the table entries consistent with the text labels.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored by many of the same researchers whose constructions supply the core definitions and representative schemes (e.g., [28], [68], [73], [75], [30]). This is not improper per se, but combined with the non-reproducible selection in Section 1.1, it makes it difficult to rule out a self-selection bias in Table 1 and Section 6. I would encourage the editor to require a full corpus list and an explicit comparison with prior surveys such as [99] regarding scope and coverage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the SoK and mostly agree with the conditional verdict. It's the closest thing we have to a shared reference for GenAI watermarking, and the Section 3 formalization is useful even though it largely restates definitions from prior work. The paper's real contributions are organizational: the clean separation of quality notions (empirical, low-distortion, distortion-free, undetectable), the detection-vs-attribution distinction in Section 3.4, and a thoughtful open-problems list in Section 7. Table 1's qualitative comparison is honest about trade-offs, and the writing is clear enough that a policymaker could follow it.\n\nThe soft spot is the selection methodology. Section 1.1 says 'over 100 peer-reviewed papers' from chosen venues, but there's no search protocol, no inclusion/exclusion rules, and no list of papers. For a survey whose central claim is comprehensiveness, that's a reproducibility gap. If the selection is biased, the taxonomy and Table 1 could mislead more than organize. I'd push the authors to either publish the full corpus and selection protocol or soften the claim to 'representative.'\n\nThe stress-test note about authorial over-weighting is fair but not fatal. Several core definitions (undetectability, distortion-freeness, unforgeability) come from the authors' own papers. Those are strong, peer-reviewed results, so citing them is reasonable. But it does mean the survey's organization reflects one research program's priorities. An independent survey might draw the boundaries differently. I'd flag it as a point for discussion, not a flaw.\n\nMinor issues: the claim in Section 3.1.3 that undetectability prevents learning detection keys is asserted without proof. The empirical evaluation section is mostly standard metrics and could be more critical about what the metrics miss. These are minor.\n\nWho's it for? Graduate students entering watermarking, and technically inclined policymakers. I'd cite it as the standard reference for definitions. Should it go to peer review? Yes. It deserves serious refereeing, with the caveat that the comprehensiveness claim needs support before publication. I'd bring it to our reading group as a useful survey.","headline":"A valuable, well-organized SoK that will be the standard reference for GenAI watermarking definitions, though its comprehensiveness claim rests on an unlisted, non-reproducible paper selection.","tokens_in":26389,"tokens_out":3105,"would_cite":true,"duration_ms":26734,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper sets out to give the field of watermarking for AI-generated content a single formal vocabulary—definitions of quality, false positive rate, robustness, unforgeability, and message support—and to map how representative schemes…","keywords":["watermarking","AI-generated content","content detection","text watermarking","image watermarking","audio and video watermarking","undetectability","unforgeability"],"falsifier":"A systematic literature search could collect all watermarking papers for generative models and check whether every scheme maps onto one of the six defined properties; a scheme that is deployed and works but satisfies none of the definitions, or a quantitative benchmark that reverses a rating in the paper's comparison table, would show the formalization is not field-wide.","tokens_in":25378,"feed_emoji":"🏷️","tokens_out":7472,"duration_ms":66490,"temperature":0.7,"pith_summary":"Watermarking embeds a hidden signal in AI-generated content so it can later be identified as machine-made. The paper argues that this is the most reliable route to distinguishing AI from human output, because passive detectors rely on statistical quirks that keep failing as models improve. Its central contribution is a formal framework: it defines a watermarking scheme by its generation, detection, decoding, and attribution algorithms, and it pins down six desired properties. It then uses that framework to survey over 100 papers across text, image, audio, and video, comparing representative schemes on quality, false positive rate, robustness, unforgeability, message support, and efficiency. If the framework holds, researchers and policymakers get a common language for what a watermark can and cannot guarantee.","feed_headline":"Watermarking AI content gets formal definitions and a unified survey","feed_subtitle":"One vocabulary spans text, image, audio, and video so researchers and regulators can compare what watermarks actually promise.","key_machinery":"The central machinery is the formal syntax of a watermark: a generation algorithm keyed by $gk$ that maps a prompt (and optional message) to a response, together with detection, decoding, and attribution algorithms with possibly separate keys. The load-bearing definitions are distortion and distortion-freeness ($\\max_{m,\\pi} \\frac{1}{2} \\sum_{x\\in R} |\\Pr[M(\\pi)\\to x]-\\Pr[\\mathrm{Watermark}^M_{gk}(m,\\pi)\\to x]|$), undetectability, false positive rate, robustness with respect to a channel, and unforgeability. These definitions do the work of turning a scattered literature into a single comparison table, since every scheme can be rated on the same axes.","core_discovery":"On the paper's own terms, the discovery is that the many ad-hoc watermarking proposals can be understood as instances of a single design space with six formal axes. The paper formalizes quality from weakest to strongest: empirical validation, low distortion, distortion-freeness where the distribution of a single watermarked response matches the unwatermarked model, and undetectability, where even adaptive queries cannot tell the watermarked model apart. It defines false positive rate against content produced independently of the key, robustness against a channel that transforms content, and unforgeability, which requires an attacker to have the embedding key to create falsely attributed content. It separates detection from attribution because the two need different guarantees: detection should be robust, attribution should be unforgeable. It also notes that text watermarks can only embed a few bits, while images and audio can carry messages, and that undetectability is impossible statistically, only computationally. The paper's contribution is this ordering and separation, not a new watermark construction.","pith_inferences":["A natural next step, not taken in the paper, is to turn the qualitative comparison table into a reproducible benchmark by running standard attack suites on each representative scheme and reporting the numbers.","The paper's warning that undetectable watermarks can be used to track individuals suggests that privacy regulation should govern who may embed and who may read watermarks; the paper flags the risk but does not develop the policy consequence.","If proof-of-generation becomes feasible, watermark keys will likely need public commitments, so that a model owner cannot retroactively pick a key that makes arbitrary content appear watermarked.","The taxonomy implies that 'watermark' is not one technology: text, image, audio, and video schemes have different entropy headroom, so deployment standards should be modality-specific rather than uniform."],"forward_implications":["If the framework is accepted, future watermarking papers can state precisely which properties they satisfy, and reviewers can ask whether a missing property matters for the stated deployment scenario.","Policymakers get a checklist: laws that mandate watermarking should specify whether they require detection, attribution, message embedding, or unforgeable public attribution, because these are different capabilities.","The separation of detection and attribution means a single generation algorithm can support both robust spam filtering and unforgeable accountability, without one property destroying the other.","The entropy caveat implies that watermarking cannot be a universal detector: near-deterministic outputs cannot carry a watermark, so regulations expecting all AI output to be marked will hit a technical limit."],"supporting_citations":[{"why":"Green-red watermarking: the core LLM logit-biasing scheme used to illustrate quality and robustness trade-offs in text.","marker":"[26]"},{"why":"Undetectable watermarks: supplies the undetectability definition and the entropy condition needed for detectable watermarks.","marker":"[28]"},{"why":"Distortion-free watermarks: grounds the formal definitions of distortion-freeness and a fixed-key robust text watermark.","marker":"[68]"},{"why":"Pseudorandom error-correcting codes: the basis for the survey's treatment of watermarks that are both undetectable and robust.","marker":"[73]"},{"why":"Unforgeable public attribution: grounds the unforgeability definition and the detection/attribution separation in Section 3.4.","marker":"[75]"},{"why":"Tree-ring watermark: a representative image watermark whose limitations the survey uses to motivate later schemes.","marker":"[29]"},{"why":"A deployed production system used as the example of large-scale watermarking with human evaluation and low false positive rates.","marker":"[57]"},{"why":"A regulatory document in Section 2.4 that motivates the watermarking requirements the framework is meant to inform.","marker":"[53]"}],"fun_headline_variants":["Formalizing AI watermarks: six axes, one framework","Watermarking AI content gets a unified design space","Six axes that define every AI watermarking scheme","AI watermarking: from empirical to undetectable","A single vocabulary for AI content watermarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the informally chosen set of over 100 papers is representative enough that the taxonomy and comparison generalize to the whole field; the paper gives no reproducible search protocol, inclusion rules, or corpus list.","fun_headline_variants_meta":{"raw":{"variants":["Formalizing AI watermarks: six axes, one framework","Watermarking AI content gets a unified design space","Six axes that define every AI watermarking scheme","AI watermarking: from empirical to undetectable","A single vocabulary for AI content watermarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000628,"raw_usage":{"total_tokens":2904,"prompt_tokens":943,"completion_tokens":1961,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1886}},"tokens_in":559,"tokens_out":1961,"duration_ms":13498,"temperature":1.0,"reasoning_tokens":1886,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:08:50.235374+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic literature search could collect all watermarking papers for generative models and check whether every scheme maps onto one of the six defined properties; a scheme that is deployed and works but satisfies none of the definitions, or a quantitative benchmark that reverses a rating in the paper's comparison table, would show the formalization is not field-wide.","supporting_citations":[],"review_version":1}