{"id":"0b0f0692-3468-4c64-84b2-4247ebb0437b","arxiv_id":"2608.06416","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"WorldMark modulates watermark strength per token using knowledge-graph saliency, improving robust detection and perplexity for MorphMark watermark variants on C4.","lead":"This paper introduces WorldMark, an add-on that uses a world-knowledge memory graph to decide where to strengthen or relax a language model watermark during text generation. In C4 experiments, WorldMark improved watermarked-text detection under synonym attacks while slightly reducing perplexity, with no changes to the detector.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline-fidelity concern: WorldMark's headline +0.0296 robust-TPR gain is measured against repro MorphMark baselines that underperform the original published values; a weak or mis-configured baseline could fully account for the gain.","rationale":"The reader's weakest assumption identifies exactly the condition the central claim depends on: the delta from internally reproduced MorphMark baselines. The gap between 'Repro' and 'Paper' values (exp 0.9000 vs 0.9600; log 0.8525 vs 0.9375) is large enough that if the reproduction is at a lower watermark-strength operating point, WorldMark's +0.0296 robust TPR gain may largely be a repair of the weak baseline. This is a validity threat to the comparison, not a disagreement with the field; the method may still work, but the evidence as reported is conditional. The ablation study (Table 9) rules out some alternatives, such as irrelevant context underperforming the full method, so the concern is specifically about host baseline fidelity. A direct run of the official MorphMark implementation settles it. I agree with the reader's CONDITIONAL verdict and do not move it; no new concern justifies acceptance or rejection beyond what the reader already stated.","tokens_in":18800,"tokens_out":12559,"duration_ms":126220,"concrete_test":"Run the official MorphMark code from arXiv:2505.11541 with its own evaluation script on the exact C4/OPT-1.3B Word-S protocol used in Table 1 (400 continuations, 30-token prefix, 200-230 tokens, seeds 0-4). If the official exp/linear/log robust TPR@1% values reproduce near 0.9600/0.9275/0.9375 while the paper's 'Repro' column reports 0.9000/0.9000/0.8525, recompute the Table 2 deltas against the official baselines; if +WorldMark (0.9119/0.9495/0.8800) no longer beats all three, the headline gain is an artifact of the weakened reproduction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is defined by deltas over internally reproduced MorphMark baselines (Table 1/2), so its truth depends on those reproductions being faithful implementations at the same operating point as the host. They are not: for exp/linear/log, the repro robust TPR@1% is 0.9000/0.9000/0.8525 versus Paper's 0.9600/0.9275/0.9375, and +WorldMark reaches 0.9119/0.9495/0.8800, still below Paper for exp and log. The repro baselines also have substantially lower PPL (e.g., exp 10.9404 vs 11.3569), consistent with a lower watermark-strength regime rather than faithful reproduction. Since AKM multiplies the host strength by mu_t (Eqs. 8-9), a low-strength reproduction can be boosted trivially. If the repro pipeline is under-strength or uses a different Word-S attack configuration, WorldMark's +0.0296 average robust TPR and +0.0108 F1 are partly artifacts of the weak baseline, not evidence that knowledge saliency improves the genuine host. The paper itself notes 'reproduced baselines do not always match' and that UW/DiPmark robust TPR drop from 0.7425/0.7250 to 0.2975/0.2700, reinforcing the risk. Without released code or a matched-operating-point comparison, the central claim remains conditional on baseline fidelity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes WorldMark, a plug-and-play interface that retrieves semantic and episodic knowledge from a World Knowledge Memory graph, computes a token-level knowledge saliency score, and adjusts the strength of a host watermark via Asymmetric Knowledge Modulation without changing the host detector. The primary experiments on C4 with OPT-1.3B compare three MorphMark host variants with and without WorldMark, reporting gains in clean and attacked detection and small perplexity reductions relative to internally reproduced MorphMark baselines. Extended evaluations on larger backbones and QA datasets, a scaled cross-family study, and a small pilot that injects memory directly without modulation are also presented. The paper includes ablations, hyperparameter tables, seed-level results, and paired t-tests.","tokens_in":19114,"tokens_out":8844,"duration_ms":83438,"significance":"If the reported effects are real, WorldMark is a useful contribution: it is detector-preserving, host-agnostic, and shows a simultaneous improvement in robustness and perplexity, a direction of genuine interest in LLM watermarking. The paper is methodologically careful in several ways: validation, threshold calibration, and test data are disjoint; five seed runs with paired t-tests are reported; and the ablations distinguish knowledge content, retrieval quality, saliency source, and modulation mechanism, including negative controls such as shuffled retrieval and random saliency. There is no obvious circularity: the detector is unchanged and the saliency score is derived from the model's own distribution. The main uncertainty is empirical: the central claim is defined as a delta over reproduced baselines that differ substantially from the originally published MorphMark values.","major_comments":[{"comment":"The headline result of an average +0.0296 robust TPR gain is a delta against the internally reproduced MorphMark baselines, and those baselines are not at the operating point of the published host watermark. In Table 1, MorphMark-exp robust TPR@1% is 0.9000 (Repro) vs 0.9600 (Paper), MorphMark-log is 0.8525 vs 0.9375, and the repro PPL values are noticeably lower (10.9404 vs 11.3569 for exp), consistent with a lower-strength regime. Because AKM multiplies the host strength by μ_t (Eqs. 8–9), a low-strength baseline can be boosted trivially, and for exp and log the +WorldMark values (0.9119 and 0.8800) still do not reach the published MorphMark values. The central claim therefore depends on these reproductions being faithful implementations at a matched operating point. Please provide either (i) a reproduction configuration whose robust TPR and PPL match the published MorphMark values, (ii) a direct comparison of +WorldMark against the published values, or (iii) released code and the exact attack/evaluation configuration so the operating point can be verified. The paper's own statement that 'reproduced baselines do not always match' (§4.2) makes this issue explicit rather than hypothetical.","section":"§4.1, Tables 1–2, Eqs. (8)–(9)"},{"comment":"Figure 2 contains numerical values that contradict Table 1, which is the paper's main results table. In panel (a), for example, MorphMark-exp is shown as 87.2/87.9/92.1 for Paper/Repro/+WorldMark, while Table 1 reports robust TPR@1% of 0.9600/0.9000/0.9119; for KGW the figure shows 46.6/45.6 whereas Table 1 reports 0.8050/0.6775. Panel (b) similarly reports values such as 59.8/59.2/66.2 for MorphMark-exp robust Best F1 while Table 1 reports 0.9778/0.9672/0.9783. The figure therefore does not illustrate the results it is intended to present. Please reconcile the figure with Table 1 or remove it.","section":"Figure 2, Tables 1"},{"comment":"The claim that every monitored metric moves in the favorable direction is overstated. For the log variant, ΔTPR@1% is 0.0000, and for exp and linear the clean TPR gains are only +0.0025 and +0.0050 on already near-saturated values (0.9975–1.0000). The abstract's statement that WorldMark 'improves clean and attacked detection' should be qualified, or the clean TPR differences should be shown to be statistically significant. The robust metrics, which show larger and significant deltas, are the appropriate focus of the paper's contribution.","section":"§4.3, Table 2, Abstract"}],"minor_comments":[{"comment":"The caption states 'All checked entries are improvements,' but the log variant's ΔTPR@1% is 0.0000; use 'non-decreasing' or qualify the statement.","section":"Table 2"},{"comment":"The extended evaluations claim to 'confirm' transfer to LLaMA-3-8B, Mistral-7B, TriviaQA, and NaturalQuestions, but these tables report only robust TPR and Robust F1 without standard deviations, confidence intervals, significance tests, or perplexity values; consider describing these as consistent exploratory evidence rather than confirmation.","section":"Appendix A, Tables 3–4"},{"comment":"The claim of no measurable latency penalty is based on wall-clock differences of about 0.01–0.16 s that are within normal runtime variation, but the mechanism is stated as GPU asynchrony; please report the isolated per-step Sentence-BERT inference time and clarify how it is hidden in batch size 1.","section":"§3.3, Table 1"},{"comment":"In Algorithm 1 the notation 'topm(p^k_t)' should be 'top_m(p^k_t)' to denote the set of top m tokens, and the loop bound T should be defined in the algorithm or by reference to the generation-length protocol.","section":"Algorithm 1, §3.3"}],"recommendation":"major_revision","confidential_remarks":"The main vulnerability is baseline fidelity: the headline numerical gains are deltas over internally reproduced MorphMark baselines that underperform the published values, and the paper itself acknowledges the mismatch. A matched-operating-point comparison or code release is needed before the central claim can be accepted. The figure/table inconsistency should also be corrected during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: genuinely novel idea — using a knowledge-saliency signal to decide where watermark strength goes — and the paper is honest about its limits, but the headline numbers are computed against reproductions that underperform the original MorphMark numbers, so the measured benefit is not yet established.\n\nWhat is actually new: existing watermark families (logits, sampling, entropy, adaptive) all place signals using local token statistics. WorldMark instead uses a semantic-episodic memory graph to compute a token-level knowledge saliency score and modulates the host watermark strength accordingly. That is a real departure. The design is also clean: no detector changes, no backbone retraining, and prompt-level retrieval so the per-token overhead is just a small embedding similarity. The paper includes a useful negative control — direct memory injection without modulation is unstable — which motivates the AKM design. The ablation study (Appendix E) is well-constructed and shows that each component contributes, and that shuffled or irrelevant context hurts. They use disjoint validation/test sets and paired t-tests over five seeds, which is more than many watermarking papers do.\n\nThe main soft spot is baseline fidelity. The reproduced MorphMark baselines are substantially below the published values: exp robust TPR 0.9000 vs 0.9600, log 0.8525 vs 0.9375, and the repro PPL is lower, consistent with a weaker watermark-strength regime. WorldMark's headline +0.0296 average robust TPR is computed against these reproductions, and for exp and log it still does not reach the original published robust TPR. The paper acknowledges the reproductions don't always match, but that makes the central claim conditional on the reproductions being faithful at the same operating point as the host. Without released code or a matched-operating-point comparison, part of the gain could be an artifact of a weak baseline.\n\nTwo smaller issues. The 'negligible overhead' claim excludes the one-time cost of running an 8B-parameter fact extractor per episode; that is a real deployment cost. And the cross-host claim rests mostly on a 12-sample pilot; the scaled evaluation in Appendix B shows AUROC/F1 gains but not under attack.\n\nThese are addressable, not fatal. The idea is worth pursuing. I would send it to peer review, but I would insist on code release, or at least a comparison that matches the original MorphMark operating point, before accepting.\n\nWho it's for: researchers working on LLM watermarking robustness and quality trade-offs. It will be most useful as a discussion piece about how to evaluate plug-and-play watermark interfaces.","headline":"Novel knowledge-saliency watermarking idea, but the headline gains rest on reproductions that underperform the original MorphMark numbers, so the central claim is conditional.","tokens_in":19663,"tokens_out":3140,"would_cite":false,"duration_ms":29522,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WorldMark claims a plug-and-play knowledge-saliency interface that improves attacked detection and text quality at the same time, across three adaptive-strength MorphMark host variants, without changing the detector or retraining the…","keywords":["LLM watermarking","knowledge saliency","world knowledge memory","asymmetric knowledge modulation","adaptive-strength watermarking","plug-and-play interface","attacked detection","C4 evaluation"],"falsifier":"Independently re-run the primary C4 evaluation with a verified faithful implementation of MorphMark that matches the original published attacked TPR values (0.9600 for exponential, 0.9275 for linear, 0.9375 for logarithmic) rather than the reproduction values; if the average attacked TPR gain over that faithful baseline is not positive, the central claim fails.","tokens_in":18531,"feed_emoji":"🧠","tokens_out":5492,"duration_ms":54478,"temperature":0.7,"pith_summary":"WorldMark proposes that where an LLM watermark is placed should depend on which tokens carry confirmed world knowledge, not just on local token probability. The paper builds a semantic-episodic memory graph from the prompt and observed context, scores each decoding position by how strongly the candidate distribution aligns with retrieved knowledge, and uses that saliency score to relax the watermark on anchored tokens while strengthening it on unanchored ones. On the primary C4 protocol with three MorphMark host variants, the complete interface reports higher clean and attacked detection and slightly lower perplexity than the reproduced baselines, with average attacked TPR +0.0296 and attacked Best F1 +0.0108. Direct memory injection without saliency-aware modulation is unstable across hosts, which the paper treats as evidence that the modulation mechanism, rather than the knowledge content alone, carries the benefit. If the pattern holds, knowledge-aware placement offers a host-agnostic way to make watermarks more attack-resistant without retraining the generator or modifying the detector.","feed_headline":"Plug-in knowledge signal lifts LLM watermark detection on three hosts","feed_subtitle":"Average attacked-detection TPR rises 0.0296 while perplexity falls, the paper reports.","key_machinery":"The load-bearing object is the knowledge saliency score $s_t$, a soft cosine similarity between the serialized retrieved knowledge context and the top-$m$ candidate tokens at each decoding step. It is converted into the asymmetric modulation factor $\\mu_t = (1-\\beta_p s_t)(1+\\beta_d(1-s_t))$, which is the mechanism that moves watermark strength away from knowledge-anchored positions and toward unanchored ones. Because the knowledge context $c_0$ is fixed per generation episode while the candidate distribution $p_t^k$ changes with the prefix, the score is position-dependent without requiring per-token retrieval or access to future tokens. The same $\\mu_t$ is applied to adaptive-logits strengths, sampling probabilities, and entropy-gate thresholds, which is what makes the interface host-agnostic while leaving the original detector untouched.","core_discovery":"The central claim is that a token-level knowledge saliency signal derived from a world-knowledge memory graph can be spliced into existing watermarks as a plug-and-play interface, improving attacked detection and text quality at the same time. Specifically, WorldMark combines World Knowledge Memory (a semantic-episodic graph with factual triples and supporting observations), a Knowledge Saliency Estimator computing $s_t = \\sigma(\\lambda \\cdot \\mathrm{sim}(\\varphi(c_0), \\varphi(\\mathrm{top}_m(p_t^k))))$, and Asymmetric Knowledge Modulation with a quality-relief coefficient $\\rho_t = 1-\\beta_p s_t$ and a detection-boost coefficient $\\eta_t = 1+\\beta_d(1-s_t)$. The composite factor $\\mu_t = \\rho_t \\eta_t$ scales the host watermark strength, and because it correlates with $1-s_t$ on high-entropy positions, it concentrates signal on weakly grounded tokens while keeping knowledge-anchored tokens close to their natural form. The paper reports consistent improvements for exponential, linear, and logarithmic MorphMark strength functions on OPT-1.3B over its own reproductions, with the linear variant showing the largest gain under synonym replacement.","pith_inferences":["If the mechanism generalizes, the same saliency-guided placement could be combined with stronger attacks such as LLM-based paraphrase or translation round-tripping: attackers would be forced to rewrite anchored factual tokens, which should make the watermark harder to erase than under uniform placement.","The ablations suggest that knowledge similarity outperforms random and entropy-based saliency, implying that semantic grounding carries information beyond local uncertainty; a direct test would swap in other knowledge sources, such as retrieval-augmented passages, to see whether the saliency signal rather than the graph structure is what matters.","Because the memory graph is built only from the prompt and already observed text, the interface could in principle be deployed in black-box settings where the watermarking party controls decoding but not the model weights or detector; a direct test would be to attach WorldMark to a hosted API generation endpoint.","The use of a small 22.7M-parameter encoder for saliency suggests that knowledge-aware placement could be added to existing watermarking pipelines with only modest auxiliary compute at generation time, though whether this overhead stays negligible on slower or larger configurations remains an open question."],"forward_implications":["A placement signal based on external knowledge rather than local entropy can be shared across watermark families that otherwise use different signal constructions.","The quality versus detectability trade-off can be eased in both directions at once: anchored tokens are perturbed less, while unanchored positions carry a stronger signal, so attacked detection rises as perplexity falls.","The host detector stays unchanged, so existing deployment and verification pipelines could adopt WorldMark by modifying only decoding-time modulation.","The paper claims the improvements carry over to LLaMA-3-8B, Mistral-7B, TriviaQA, and NaturalQuestions, and that the benefit widens on knowledge-intensive datasets where entity-bearing tokens carry greater semantic weight."],"supporting_citations":[{"why":"Defines the MorphMark adaptive-strength host watermark family that WorldMark attaches to and whose reported results serve as the reproduction baselines.","marker":"[8]"},{"why":"Supplies the semantic-episodic memory graph organization used as World Knowledge Memory.","marker":"[9]"},{"why":"Provides the green-red logits watermarking scheme underlying the adaptive-logits instantiation and the observation that paraphrasing weakens watermarks.","marker":"[1]"},{"why":"Provides the fact-extraction prompt template used to build semantic triples for the memory graph.","marker":"[27]"},{"why":"Provides the sentence embedding encoder used to compute knowledge-saliency similarity scores.","marker":"[28]"}],"fun_headline_variants":["WorldMark: Knowledge saliency boost for watermark detection","Plug-and-play knowledge graph lifts LLM watermark detection","Memory-graph interface strengthens watermark detection and quality","Knowledge-aware watermarking cuts perplexity, boosts detection","Saliency-tuned watermarking improves robustness without retraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the internally reproduced MorphMark baselines are faithful implementations of the host watermarks; the reproduction numbers fall below the original published values, so the reported gains could shrink if those baselines are too weak.","fun_headline_variants_meta":{"raw":{"variants":["WorldMark: Knowledge saliency boost for watermark detection","Plug-and-play knowledge graph lifts LLM watermark detection","Memory-graph interface strengthens watermark detection and quality","Knowledge-aware watermarking cuts perplexity, boosts detection","Saliency-tuned watermarking improves robustness without retraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1336,"prompt_tokens":999,"completion_tokens":337,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":261}},"tokens_in":615,"tokens_out":337,"duration_ms":3884,"temperature":1.0,"reasoning_tokens":261,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:26:01.449246+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently re-run the primary C4 evaluation with a verified faithful implementation of MorphMark that matches the original published attacked TPR values (0.9600 for exponential, 0.9275 for linear, 0.9375 for logarithmic) rather than the reproduction values; if the average attacked TPR gain over that faithful baseline is not positive, the central claim fails.","supporting_citations":[{"cited_title":"Sorokin, and Mikhail Burtsev","cited_arxiv_id":null,"evidence_quote":"Provides the fact-extraction prompt template used to build semantic triples for the memory graph."},{"cited_title":"Sentence-bert: Sentence embeddings using siamese bert-networks","cited_arxiv_id":null,"evidence_quote":"Provides the sentence embedding encoder used to compute knowledge-saliency similarity scores."}],"review_version":1}