REVIEW 2 major objections
Attributing a watermarked text to one of N users costs Θ(log N/h) tokens over every stationary-ergodic source of entropy rate h, sharp to a (1+o(1)) factor.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 01:38 UTC pith:KKDE6HIY
load-bearing objection Abstract-only: a sharp, two-sided Θ(log N/h) multi-user attribution law for distortion-free watermarks looks like a real first, but proofs and experiments are invisible. the 2 major comments →
Watermark Forensics for Generative Models: An Information-Theoretic Perspective
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
For statistically distortion-free watermarks, multi-user attribution among N identities requires Θ(log N/h) tokens on every stationary-ergodic source of entropy rate h, sharp to a (1+o(1)) factor, and is the first tight entropy-rate law for that task; extraction of an ℓ-bit payload likewise costs Θ(ℓ/h). The achievability side uses exact alignment with a per-candidate surprisal decoder that almost never implicates an innocent user.
What carries the argument
The information profile ν(t) = I(S; X_t | X_<t). Its total mass determines the sample length needed for attribution and extraction; how that mass is spread determines localization cost; and its two natural caps (subtle-on-every-token versus loud-on-few-tokens) recover the literature's two quality models.
Load-bearing premise
The watermark is statistically distortion-free (marked and unmarked token distributions are identical) and the source is stationary-ergodic with a well-defined entropy rate h; the matching upper bound further needs a decoder that thresholds each candidate by its own realized surprisal.
What would settle it
On any stationary-ergodic source of known entropy rate h, measure the shortest watermarked length at which a multi-user decoder correctly attributes among N identities with vanishing false-imputation probability; if that length is not asymptotically (1+o(1)) log N / h, or if collision counting matches the surprisal decoder's rate without overcharging, the claimed law fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript studies the sample-length cost of forensic tasks for watermarks in generative-model outputs (detection, multi-user attribution, payload extraction, and localization). It organizes these tasks via an information profile ν(t)=I(S;X_t|X_<t) whose total mass pays for attribution and extraction and whose spread pays for localization; detection is paid for by presence (distributional distance) rather than mutual information. The main claimed result is a sharp two-sided Θ(log N/h) law for attributing a text to one of N users under statistically distortion-free schemes, holding for every stationary-ergodic source of entropy rate h and attained only by a per-candidate surprisal-threshold decoder (not collision counting); extraction of an ℓ-bit payload costs Θ(ℓ/h). Two further gaps (a Θ(log N)-token unattributable window and a footprint-resolution uncertainty principle) are asserted, and experiments on GPT-2, Pythia-410M, and Qwen2.5 are said to recover the predicted constants.
Significance. If the claimed (1+o(1))-sharp entropy-rate law and the matching converse hold as stated, the paper would supply the first tight multi-user attribution length for statistically distortion-free watermarks on stationary-ergodic sources, together with a concrete decoder that avoids the unbounded overcharge of collision counting. That would be a genuine advance for the information-theoretic foundations of watermark forensics and would clarify the separation between detection and attribution. The information-profile framing and the reported recovery of constants on three language models would further strengthen the contribution, provided the proofs and experimental protocols are complete and reproducible.
major comments (2)
- Only the abstract is available for review. The central (1+o(1))-sharp Θ(log N/h) attribution law, the matching converse, the surprisal-threshold decoder construction, the claimed failure of collision counting, the two asserted gaps, and the experimental recovery of constants on GPT-2/Pythia/Qwen cannot be verified. Without the full text, theorems, proofs, and experimental details, it is impossible to confirm that the load-bearing claims hold or that the decoder attains the stated rate while controlling false implication of innocents. A full manuscript is required before any technical assessment of correctness can be made.
- The abstract asserts that the law holds for every stationary-ergodic source of entropy rate h under statistically distortion-free schemes. The precise statement of the distortion-free constraint, the regularity conditions on the source class, and the exact alignment notion used for multi-user attribution are not visible. These definitions are load-bearing for both the achievability and the converse; they must be supplied and checked for internal consistency with the information-profile mass bound ν(t) ≤ h.
Circularity Check
No significant circularity visible in the abstract; rates follow standard mutual-information accounting under stated assumptions.
full rationale
Only the abstract is available, so the full derivation chain cannot be inspected equation-by-equation. On the face of the abstract the central claims are organized by an information profile ν(t)=I(S;X_t|X_<t) whose total mass must accumulate to log N (or ℓ) for attribution (or extraction). The statistically distortion-free constraint forces that mass to be paid at rate at most the source entropy rate h, yielding the Θ(log N/h) and Θ(ℓ/h) scalings; a matching converse and an explicit per-candidate surprisal-threshold decoder are asserted to close the (1+o(1)) factor. These are standard information-theoretic lower/upper bounds relative to an external source class (stationary-ergodic of entropy rate h), not quantities defined in terms of the target rates themselves, not fitted parameters renamed as predictions, and not uniqueness theorems imported solely by self-citation. No self-definitional loop, fitted-input-called-prediction, or ansatz-smuggling step is quotable from the abstract. Residual risk is purely that the unavailable full proofs might later introduce circular steps; under the given text the honest finding is score 0 with empty steps.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Source is stationary-ergodic with entropy rate h
- domain assumption Watermark schemes are statistically distortion-free (marked and unmarked distributions identical)
- standard math Standard mutual-information and asymptotic equipartition machinery
invented entities (1)
-
information profile ν(t)=I(S;X_t|X_<t)
no independent evidence
read the original abstract
A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length $n$. One object organizes the answers. Let $S$ be the secret the mark carries (a user's identity or payload), and let the information profile $\nu(t)=I(S;X_t\mid X_{<t})$ record how much the $t$-th token reveals about $S$ given the earlier ones. Its total mass pays for attribution and extraction; how that mass is spread pays for localization; and detection alone is paid for not by information but by presence, the distance from the marked to the unmarked distribution. The literature's two quality models, a mark subtle on every token and one that stamps a few tokens loudly, are two incomparable ways of capping this profile. Our main theorem settles the ladder's entropy column. For statistically distortion-free schemes, attributing a text to one of $N$ users costs $\Theta(\log N/h)$ tokens over every stationary-ergodic source of entropy rate $h$, sharp to a $(1+o(1))$ factor: to our knowledge the first tight entropy-rate law for multi-user attribution (via exact alignment). The natural collision-counting analysis overcharges without bound; only a decoder thresholding each candidate by its own realized surprisal attains the rate while almost never implicating an innocent user. A matching converse makes the law two-sided, and extraction of an $\ell$-bit payload costs $\Theta(\ell/h)$. Two gaps are real, not modeling artifacts: a $\Theta(\log N)$-token window in which a text is provably machine-made yet unattributable, and a footprint-resolution uncertainty principle. Experiments on GPT-2, Pythia-410M, and Qwen2.5 recover the predicted constants.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.