Pith. sign in

REVIEW 3 major objections 5 minor 54 references

PathMark turns MoE routing into a secret path that proves model ownership without wrecking utility.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 00:40 UTC pith:ZGQT3QA4

load-bearing objection First MoE-native path watermark that actually works; empirics are strong, stealth theory has a known α gap that does not kill the result. the 3 major comments →

arxiv 2607.03688 v1 pith:ZGQT3QA4 submitted 2026-07-04 cs.CR

PathMark: Protecting Intellectual Property of Mixture-of-Expert LLMs via Path Watermarks

classification cs.CR
keywords Mixture-of-Expertsmodel watermarkingintellectual property protectionrouting pathdistribution alignmentcontrastive lossLLM security
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Dense-model watermarks fail on Mixture-of-Experts language models because routing is dynamic: watermarked parameters are not always active, decision margins are razor-thin, and gradients concentrate on the few selected experts, erasing the mark. PathMark flips the problem into a feature. It steers every token of a secret trigger through a predetermined subset of experts in selected layers, writing a distinctive multi-expert path signature. Three design pieces make the signature stick: a distribution-alignment loss that lifts target experts to near-certainty, a wide-path rule that accepts any of several experts so partial removal fails, and a contrastive loss that cancels gradient spillover so clean inputs keep their natural routes. The same combinatorial choice of experts also encodes multi-bit identifiers. Ownership can be checked either by inspecting the router (white-box) or by a lightweight trigger-to-mark mapping (black-box). On four production-scale MoE models the method reports near-perfect detection with less than two percent perplexity rise and holds up under quantization, fine-tuning, pruning and adaptive routing attacks.

Core claim

PathMark is the first watermark built for MoE architectures: by actively constraining triggered tokens onto predetermined multi-expert routing paths via distribution-alignment and contrastive losses, it achieves greater than 99 percent verification accuracy with less than 2 percent perplexity degradation and remains detectable after quantization, fine-tuning, pruning and adaptive routing attacks across four MoE models.

What carries the argument

The dual-objective routing loss: a distribution-alignment term that forces triggered tokens onto a wide target-expert set, plus a contrastive separation term that cancels gradient leakage so clean tokens keep their original routes.

Load-bearing premise

A single fixed contrastive weight is assumed to cancel gradient spillover onto clean tokens so that clean routing stays statistically indistinguishable from the original model, even though the paper's own derivation and the numbers used in training do not fully match.

What would settle it

Train PathMark on a standard MoE, then measure the KL divergence (or expert-selection histograms) of clean inputs against the unwatermarked baseline after the claimed contrastive weight; if the distributions diverge enough for an adversary to detect the watermark with a few thousand clean queries, the stealth claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. PathMark proposes the first watermarking scheme designed for Mixture-of-Experts LLMs. It treats expert routing as a covert channel: a secret trigger forces tokens through predetermined multi-expert “wide paths” in selected late layers, producing a verifiable routing signature. Embedding uses a dual-objective loss (distribution alignment of triggered tokens plus a contrastive term claimed to cancel gradient spillover onto clean tokens) together with standard language-modeling loss. The scheme supports multi-bit combinatorial capacity, white-box routing inspection, and an optional black-box verification mark. Experiments on four MoE models report near-100 % watermark success rates, <2 % perplexity degradation, and strong robustness under quantization, fine-tuning, pruning, and several adaptive routing attacks.

Significance. If the claims hold, PathMark fills a genuine gap: existing parameter-centric watermarks assume static dense graphs and fail under MoE’s dynamic sparse routing. The empirical package is substantial—four models, multiple attack surfaces, multi-watermark coexistence, and dual verification modes—and the combinatorial capacity formula is clean. The work therefore has clear practical value for IP protection of high-value MoE checkpoints. The theoretical apparatus (gradient-cancellation lemma, KL indistinguishability bound, Hoeffding verification bound) is a welcome attempt at formal guarantees, even though the stealthiness argument currently rests on an imperfectly calibrated free parameter.

major comments (3)
  1. Lemma D.1 (Appendix D) derives a contrastive weight α ≈ 7.1 under the paper’s own Taylor / isotropic-curvature / static-partition-function approximations, yet every experiment and the implementation use α ∈ [1, 3]. The authors acknowledge the gap but treat it as a harmless scaling issue. Because Theorem 4.1 / D.2 (clean-input KL bound) and the stealthiness claim rest on this cancellation, the manuscript needs either (i) a corrected derivation that recovers the operating range of α, or (ii) a quantitative ablation that measures residual clean-token gradient magnitude and KL divergence at the actual α used. Without that, the “provably cancels gradient leakage” claim is not load-bearing.
  2. The same stealthiness bound further relies on a hand-chosen distribution-change constant B ≈ 0.02 (Assumption D.1). No measurement of B on the trained models is reported, nor is sensitivity of the sample-complexity claim (Corollary D.2) to larger B examined. A short empirical estimate of ||g^(w) − g^(0)||₂ on clean data would make the bound falsifiable rather than free.
  3. Black-box verification (Sec. 4.4.1) requires a second fine-tuning stage that maps the trigger alone to a verification mark. While the paper correctly notes that the routing signature itself remains the primary ownership proof, the reported 100 % black-box match rates are obtained only after this extra stage. The manuscript should clarify how much of the black-box success is attributable to the routing path versus ordinary backdoor-style output fine-tuning, ideally with an ablation that freezes the router after the first stage.
minor comments (5)
  1. Figure 1(a) caption and text claim top-k probabilities “hover around ∼0.04”; the plotted density peaks near 0.04 for top-k but the top-1 mean is given as 0.21—clarify which statistic is being emphasized.
  2. Eq. (7) writes L_stealth while the surrounding text and Eq. (6) use L_contrast; unify the notation.
  3. Table 1 reports p-values under different sample sizes and success criteria for each baseline; a short note on how p-values were standardized would improve comparability.
  4. The latency-based verification sketch (Sec. 6.2) is interesting but exploratory; either move it to an appendix or supply the multi-GPU measurement protocol and variance.
  5. A few typos: “via a binomial test” (missing subject), “L (l)_stealth” vs. contrast, and “id=[1010]” in Figure 2 legend.

Circularity Check

1 steps flagged

Mild self-consistency in the stealthiness theory: the KL bound and gradient-cancellation lemma largely restate the design of the dual loss plus a hand-chosen B, with an acknowledged α theory–practice gap; empirical claims remain non-circular.

specific steps
  1. self definitional [Lemma D.1 (Appendix D) + Theorem 4.1 / Sec. 4.5.2]
    "a contrastive loss provably cancels gradient leakage to clean inputs, maintaining their natural routing path. … Under Lemma D.1, for any clean input x∼D_clean … E[D_KL(g^{(0)}_l(x)∥g^{(w)}_l(x))]≤B^{2}/p_eff_min … where … B is the maximum ℓ_{2}-norm change in routing distribution induced by watermark training."

    The dual loss is constructed so that L_contrast opposes L_align spillover; Lemma D.1 then derives cancellation for a specific α under approximations that the paper itself later concedes do not match the operating α∈[1,3]. The subsequent KL bound simply inserts that cancellation plus a hand-chosen small B justified by the same objective. The “guarantee” is therefore largely a restatement of the design assumptions rather than an independent derivation.

full rationale

PathMark’s core empirical results (WSR >99 %, PPL degradation <2 %, robustness tables) are obtained by training the dual-objective loss on mixed batches and evaluating on held-out triggers/clean data; success rates are not definitional restatements of the loss. Capacity Cap = |L_w|·log₂(N/k) is ordinary combinatorial counting. The only circularity-adjacent material is the stealthiness argument (Lemma D.1 + Theorem 4.1/D.2): the contrastive term is introduced precisely to cancel alignment spillover, the lemma then “proves” cancellation under Taylor/isotropic/static-partition approximations that themselves produce α≈7.1 while experiments use α∈[1,3], and the KL bound further inserts an assumed small B justified by the same objective. The authors note the discrepancy and treat it as a scaling-law issue. This is mild self-consistency of a supporting theoretical claim, not a load-bearing circular derivation of the main results; hence score 2 rather than 0 or higher.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 2 invented entities

The central empirical claim rests on standard MoE routing math plus several design choices and free hyperparameters (path width, layer set, loss weights, trigger form) that are selected rather than derived. Stealthiness and effectiveness theorems further rest on gradient-cancellation and bounded-distribution-change assumptions that are only approximately justified.

free parameters (6)
  • routing loss weight λ = 1.0
    Set to 1.0; sensitivity table shows WSR stable for λ in [0.1,5.0] but value is hand-chosen.
  • contrastive weight α = 3 (empirical)
    Theory predicts ~7.1; practice uses coefficient 3 (or range [1,3]). Directly controls claimed gradient cancellation.
  • path width k_l and |L_w| = k_l=2, |L_w|=6
    Default k_l=2 target experts and final 6 layers chosen via ablation for WSR/PPL trade-off.
  • verification threshold γ = 0.8
    White-box Acc≥0.8 declared watermarked; threshold is design choice.
  • distribution-change bound B = ≈0.02 (illustrative)
    Appears in Theorem 4.1 / Assumption D.1 as a small constant (~0.02 in remarks) used to bound clean-input KL; not measured as a free fit to all data but assumed.
  • contrastive temperature τ_T = 1.0
    Fixed at 1.0 in implementation; enters α formula.
axioms (5)
  • domain assumption MoE top-k routing is a controllable, sufficiently expressive channel that can be steered by router fine-tuning without destroying language-model utility.
    Foundational premise of Sec. 4; supported empirically but not proved for arbitrary MoE scales.
  • ad hoc to paper Lemma D.1 gradient balance: with suitable α, alignment leakage on clean tokens is cancelled by contrastive negative sampling so net clean routing gradient ≈0.
    Load-bearing for stealthiness theorems; derivation uses first-order Taylor, isotropic Hessian, and idealized batch assumptions that the paper itself notes do not match empirical α.
  • domain assumption Adversary does not know the secret trigger and cannot train a comparable MoE from scratch.
    Threat model Sec. 3.1; standard watermarking assumption.
  • standard math Hoeffding bound under uniform null expert selection for white-box p-values.
    Theorem 4.2 / D.3; conservative null, standard concentration inequality.
  • domain assumption Watermarking only late layers preserves hierarchical feature extraction and limits fidelity loss.
    Sec. 4.2 design rationale; common LLM lore, not independently proved here.
invented entities (2)
  • Path watermark Π* (predetermined multi-expert routing signature across L_w layers) independent evidence
    purpose: Ownership identifier realized as combinatorial expert subsets rather than static weight patterns.
    Core construct of Definition 2; independent evidence is the measurable Acc and black-box mark rates under the stated protocols.
  • Dual-objective routing loss (L_align + α L_contrast) with claimed gradient cancellation independent evidence
    purpose: Simultaneously enforce triggered path and preserve clean routing.
    Training mechanism invented for this paper; falsifiable via clean-vs-triggered expert histograms and ablation without L_contrast.

pith-pipeline@v1.1.0-grok45 · 37027 in / 3650 out tokens · 30847 ms · 2026-07-12T00:40:01.894657+00:00 · methodology

0 comments
read the original abstract

Mixture-of-Experts (MoE) large language models represent high-value intellectual property, yet existing watermarking schemes designed for dense models fail on MoE architectures due to architectural mismatch: traditional methods assume watermarked parameters are consistently activated, but MoE's dynamic routing breaks this assumption. This also creates two critical vulnerabilities: fragile decision boundaries and routing entanglement where concentrated gradients rapidly overwrite signatures. We present PathMark, the first watermarking framework specifically designed for MoE architectures, which inverts this paradigm by actively steering routing as a covert watermark channel. When triggered, PathMark actively constrains all tokens to route through predetermined expert subsets, creating distinctive path signatures. Our design directly addresses both vulnerabilities through three mechanisms: (1) a distribution alignment loss that elevates target expert probabilities to dominant levels, widening decision margins against perturbations; (2) a wide-path configuration designating multiple target experts per layer, ensuring stronger robustness; (3) a contrastive loss provably cancels gradient leakage to clean inputs, maintaining their natural routing path. Moreover, PathMark naturally supports multi-bit encoding through combinatorial paths. Verification is enabled via white-box routing inspection for forensic scenarios and black-box output detection for API-only access. Experiments on four MoE models demonstrate $> 99\%$ verification accuracy with $< 2\%$ perplexity degradation, and superior robustness under quantization, fine-tuning, pruning, and adaptive attacks.

Figures

Figures reproduced from arXiv: 2607.03688 by Linghan Chen, Qingyue Wang, Ruixuan Huang, Shuai Wang, Yuanyuan Yuan, Yudong Gao, Zimo Ji.

Figure 1
Figure 1. Figure 1: (a) Fragile decision boundaries: distribution of top [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of PathMark’s dual-objective strategy (simplified with top- [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Perplexity comparison [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Expert selection distribution comparison between [PITH_FULL_IMAGE:figures/full_fig_p010_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: Watermark robustness under fine-pruning attack. [PITH_FULL_IMAGE:figures/full_fig_p011_6.png] view at source ↗
Figure 9
Figure 9. Figure 9: Multi-watermark coexistence analysis [PITH_FULL_IMAGE:figures/full_fig_p012_9.png] view at source ↗
Figure 8
Figure 8. Figure 8: Robustness under adaptive model unlearning. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗
Figure 10
Figure 10. Figure 10: Joint ablation on path width 𝑘𝑙 and layer count |𝐿𝑤 |. 0 10 20 30 40 50 60 Expert ID 0.0 0.2 0.4 0.6 Probability Layer 20 Clean Model Watermarked Model (w/o Loss) Watermarked Model (w/ Loss) [PITH_FULL_IMAGE:figures/full_fig_p016_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Expert selection distributions in layer 20 under [PITH_FULL_IMAGE:figures/full_fig_p016_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: PathMark’s deployment with dual watermarks. Two independent watermarks embedded in Qwen1.5-MoE using [PITH_FULL_IMAGE:figures/full_fig_p017_12.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

54 extracted references · 15 linked inside Pith

  1. [1]

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauff- mann, et al. 2024. Phi-4 technical report.arXiv preprint arXiv:2412.08905(2024)

  2. [2]

    Jiacheng Cai, Jiahao Yu, Yangguang Shao, Yuhang Wu, and Xinyu Xing. 2025. UTF: Under-trained Tokens as Fingerprints——A Novel Approach to LLM Identification. InIn Proc. of the Workshop on LLM Security. 1–6

  3. [3]

    Zitian Chen, Yikang Shen, Mingyu Ding, Zhenfang Chen, Hengshuang Zhao, Erik G Learned-Miller, and Chuang Gan. 2023. Mod-squad: Designing mixtures of experts as modular multi-task learners. InIn Proc. of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 11828–11837

  4. [4]

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168(2021)

  5. [5]

    Damai Dai, Chengqi Deng, Chenggang Zhao, RX Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Yu Wu, et al. 2024. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models.arXiv preprint arXiv:2401.06066(2024)

  6. [6]

    Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al

  7. [7]

    InIn Proc

    Glam: Efficient scaling of language models with mixture-of-experts. InIn Proc. of the International conference on machine learning. 5547–5569

  8. [8]

    William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research23, 120 (2022), 1–39

  9. [9]

    Yudong Gao, Honglong Chen, Peng Sun, Junjian Li, Anqing Zhang, Zhibo Wang, and Weifeng Liu. 2024. A dual stealthy backdoor: From both spatial and frequency perspectives. InProceedings of the AAAI conference on artificial intelligence, Vol. 38. 1851–1859

  10. [10]

    Chenchen Gu, Xiang Lisa Li, Percy Liang, and Tatsunori Hashimoto. 2024. On the Learnability of Watermarks for Language Models. InIn Proc. of the International Conference on Learning Representations

  11. [11]

    Song Han, Jeff Pool, John Tran, and William Dally. 2015. Learning both weights and connections for efficient neural network.Advances in neural information processing systems28 (2015)

  12. [12]

    Zhiwei He, Binglin Zhou, Hongkun Hao, Aiwei Liu, Xing Wang, Zhaopeng Tu, Zhuosheng Zhang, and Rui Wang. 2024. Can Watermarks Survive Translation? On the Cross-lingual Consistency of Text Watermark for Large Language Models. InIn Proc. of the Annual Meeting of the Association for Computational Linguistics

  13. [13]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Under- standing. InIn Proc. of the International Conference on Learning Representations

  14. [14]

    Zhengmian Hu, Lichang Chen, Xidong Wu, Yihan Wu, Hongyang Zhang, and Heng Huang. 2024. Unbiased Watermark for Large Language Models. InIn Proc. of the International Conference on Learning Representations

  15. [15]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al . 2024. Mixtral of experts.arXiv preprint arXiv:2401.04088(2024)

  16. [16]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. InIn Proc. of the International Conference on Machine Learning. 17061–17084

  17. [17]

    John Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu, Khalid Saifullah, Kezhi Kong, Kasun Fernando, Aniruddha Saha, Micah Goldblum, and Tom Goldstein. [n. d.]. On the Reliability of Watermarks for Large Language Models. InIn Proc. of the International Conference on Learning Representations

  18. [18]

    Torsten Krauß, Jasper Stang, and Alexandra Dmitrienko. 2024. {ClearStamp}: A {Human-Visible} and Robust {Model-Ownership} Proof based on Transposed Model Training. InIn Proc. of the USENIX Security Symposium. 5269–5286

  19. [19]

    Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020. Gshard: Scaling giant models with conditional computation and automatic sharding.arXiv preprint arXiv:2006.16668(2020)

  20. [20]

    Junzhuo Li, Bo Wang, Xiuze Zhou, and Xuming Hu. 2025. Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adap- tation. InIn Proc. of the Conference on Empirical Methods in Natural Language Processing. 18489–18504

  21. [21]

    Peixuan Li, Pengzhou Cheng, Fangqi Li, Wei Du, Haodong Zhao, and Gongshen Liu. 2023. Plmmark: a secure and robust black-box watermarking framework for pre-trained language models. InIn Proc. of the AAAI Conference on Artificial Intelligence, Vol. 37. 14991–14999. CCS ’26, November 15–19, 2026, The Hague, The Netherlands Yudong Gao et al

  22. [22]

    Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun. 2024. Back- doorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models.arXiv preprint arXiv:2408.12798(2024)

  23. [23]

    Zongjie Li, Chaozheng Wang, Shuai Wang, and Cuiyun Gao. 2023. Protect- ing intellectual property of large language model-based code generation apis via watermarks. InIn Proc. of the ACM SIGSAC Conference on Computer and Communications Security. 2336–2350

  24. [24]

    Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Cheng- gang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437(2024)

  25. [25]

    Aiwei Liu, Leyi Pan, Xuming Hu, Shiao Meng, and Lijie Wen. 2024. A Seman- tic Invariant Robust Watermark for Large Language Models. InIn Proc. of the International Conference on Learning Representations

  26. [26]

    Yepeng Liu and Yuheng Bu. 2024. Adaptive Text Watermark for Large Language Models. InIn Proc. of the International Conference on Machine Learning

  27. [27]

    Yijian Lu, Aiwei Liu, Dianzhi Yu, Jingjing Li, and Irwin King. 2024. An Entropy- based Text Watermarking Detection Method. InIn Proc. of the Annual Meeting of the Association for Computational Linguistics. 11724–11735

  28. [28]

    Mitch Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993. Building a large annotated corpus of English: The Penn Treebank.Computational linguistics 19, 2 (1993), 313–330

  29. [29]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. InIn Proc. of the International Conference on Learning Representations

  30. [30]

    Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Se- won Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, et al. 2024. Ol- moe: Open mixture-of-experts language models.arXiv preprint arXiv:2409.02060 (2024)

  31. [31]

    Julien Piet, Chawin Sitawarin, Vivian Fang, Norman Mu, and David Wagner. 2025. MARKMyWORDS: Analyzing and Evaluating Language Model Watermarks. In In Proc. of the IEEE Conference on Secure and Trustworthy Machine Learning. IEEE, 68–91

  32. [32]

    Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023. Fine-tuning aligned language models compromises safety, even when users do not intend to!arXiv preprint arXiv:2310.03693(2023)

  33. [33]

    Jie Ren, Han Xu, Yiding Liu, Yingqian Cui, Shuaiqiang Wang, Dawei Yin, and Jiliang Tang. 2024. A robust semantics-based watermark for large language model against paraphrasing. InIn Proc. of the Conference of the North American Chapter of the Association for Computational Linguistics. 613–625

  34. [34]

    Shuo Shao, Yiming Li, Hongwei Yao, Yiling He, Zhan Qin, and Kui Ren. 2025. Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution. InIn Proc. of the Network and Distributed System Security Symposium

  35. [35]

    Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. InIn Proc. of the International Conference on Learning Representations

  36. [36]

    Zheyue Tan, Zhiyuan Li, Tao Yuan, Dong Zhou, Weilin Liu, Yueqing Zhuang, Yadong Li, Guowei Niu, Cheng Qin, Zhuyu Yao, et al . 2025. ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts.arXiv preprint arXiv:2510.17483(2025)

  37. [37]

    Florian Tramèr, Fan Zhang, Ari Juels, Michael K Reiter, and Thomas Ristenpart

  38. [38]

    InIn Proc

    Stealing machine learning models via prediction {APIs}. InIn Proc. of the USENIX Security Symposium. 601–618

  39. [39]

    Qingyue Wang, Qi Pang, Xixun Lin, Shuai Wang, and Daoyuan Wu. 2025. Bad- MoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts.arXiv preprint arXiv:2504.18598(2025)

  40. [40]

    Wenhan Wu, Huanghuang Liang, Jingling Yuan, Jiawei Jiang, Kanye Ye Wang, Chuang Hu, Xiaobo Zhou, and Dazhao Cheng. 2025. Zero-shot federated un- learning via transforming from data-dependent to personalized model-centric. InProceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI-25, Montreal, QC, Canada. 29–31

  41. [41]

    Jiashu Xu, Fei Wang, Mingyu Ma, Pang Wei Koh, Chaowei Xiao, and Muhao Chen. 2024. Instructional fingerprinting of large language models. InIn Proc. of the Conference of the North American Chapter of the Association for Computational Linguistics. 3277–3306

  42. [42]

    Yijie Xu, Aiwei Liu, Xuming Hu, Lijie Wen, and Hui Xiong. 2025. Mark your llm: Detecting the misuse of open-source large language models via watermarking. arXiv preprint arXiv:2503.04636(2025)

  43. [43]

    Shojiro Yamabe, Tsubasa Takahashi, Futa Waseda, and Koki Wataoka. 2024. MergePrint: Robust Fingerprinting against Merging Large Language Models. arXiv e-prints(2024), arXiv–2410

  44. [44]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  45. [45]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng He, Junyang Lin, Kai Dang, Keming Lu, Keqin Chen, Kexin Yang, Mei Li, Mingfen...

  46. [46]

    Do-hyeon Yoon, Minsoo Chun, Thomas Allen, Hans Müller, Min Wang, and Rajesh Sharma. 2025. Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!arXiv preprint arXiv:2507.03014(2025)

  47. [47]

    Jimiao Yu, Honglong Chen, Junjian Li, Linghan Chen, Yudong Gao, Weifeng Liu, and Lei Zhang. 2025. Black-Box Adversarial Defense Based on Image Decompo- sition and Reconstruction.IEEE Transactions on Multimedia(2025)

  48. [48]

    Boyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu, Chenghu Zhou, Xinbing Wang, Yu Yu, and Zhouhan Lin. 2024. Huref: Human-readable fingerprint for large language models.Advances in Neural Information Processing Systems37 (2024), 126332–126362

  49. [49]

    Jialong Zhang, Zhongshu Gu, Jiyong Jang, Hui Wu, Marc Ph Stoecklin, Heqing Huang, and Ian Molloy. 2018. Protecting intellectual property of deep neural networks with watermarking. InIn Proc. of the Asia conference on computer and communications security. 159–172

  50. [50]

    Jie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang, Yong Liu, Yu Qiao, and Jing Shao. 2025. REEF: Representation Encoding Fingerprints for Large Language Models. InIn Proc. of the International Conference on Learning Representations

  51. [51]

    Hao Zhao, Zihan Qiu, Huijia Wu, Zili Wang, Zhaofeng He, and Jie Fu. 2024. HyperMoE: Towards Better Mixture of Experts via Transferring Among Experts. InIn Proc. of the Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 10605–10618

  52. [52]

    @@@@", "!@#¥%

    Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022. St-moe: Designing stable and transfer- able sparse expert models.arXiv preprint arXiv:2202.08906(2022). A Implementation Details We watermark the final 4 to 6 MoE layers of the model and adopt a wide-path configuration with 𝑘𝑙 = 2target experts...

  53. [53]

    Increasing layer count mitigates this: at𝑘𝑙 = 10, false positives drop from 35% (1 layer) to 9% (10 layers) as multi-layer constraints become more discriminative

    produces 35% false positives because selecting many experts overlaps substantially with natural routing. Increasing layer count mitigates this: at𝑘𝑙 = 10, false positives drop from 35% (1 layer) to 9% (10 layers) as multi-layer constraints become more discriminative. (3) Perplexity scales with watermark coverage.Both 𝑘𝑙 and |𝐿𝑤 | increase perplexity, rang...

  54. [54]

    leakage" denotes the indirect effect on clean samples via shared router parameters, which also means spillover, and

    achieves 13.0 PPL. Recommendation:Configuration (𝑘𝑙 = 2,|𝐿 𝑤 |= 6) provides optimal balance: 93% WSR, 2% false positives, 13.0 PPL. B.2 Hyperparameter Sensitivity (𝜆) We analyze the system’s sensitivity to the routing loss weight 𝜆, which governs the trade-off between watermark strength and lan- guage modeling fidelity. Table 9 reveals that PathMark exhib...