REVIEW 4 major objections 4 minor 44 references
Watermarking LLMs for traceability can silently corrupt clinical reasoning—fabricating diagnoses, misusing terms, and misreading images—even when benchmark accuracy barely moves.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 13:48 UTC pith:EE32JH4D
load-bearing objection Solid empirical evidence that watermarking degrades clinical reasoning even when accuracy benchmarks stay flat; the judge calibration is the main caveat, not a fatal one. the 4 major comments →
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that applying a text watermark to a medical LLM can degrade the clinical quality of the output even when the model's benchmark accuracy remains statistically unchanged. Across five watermarking schemes, eleven language models, and seven vision-language models, the authors find that watermarks increase the rate of correct answers backed by flawed reasoning, inject fabricated or misapplied medical terms—'pentaprazole' for pantoprazole, a non-existent antibiotic 'naftifloxacin'—and shift visual grounding so that supported image findings fall while contradicted claims rise. The damage is scheme-specific and model-specific: one scheme mostly corrupts surface spelling,
What carries the argument
The load-bearing instrument is a set of three physician-validated LLM-as-judge audits that look past the multiple-choice letter. A single-response audit scores each completion for fabricated entities, misapplied real terms, corrupted spellings, confident errors, clue misreads, contradictions, and—for multimodal inputs—classifies each perceptual image claim as supported, contradicted, or unverifiable against the image. A pairwise divergence audit compares watermarked and unwatermarked answers to the same question and flags mutually exclusive diagnostic interpretations even when both pick the correct letter. A reasoning-trace audit rates the hidden chain of thought of reasoning models on effic
Load-bearing premise
The entire harm estimate rests on the LLM judge's flags being a valid proxy for clinically meaningful degradation; physician agreement is strong for fabricated entities but only moderate for misapplied terms and image-contradicted claims, so if those flags carry systematic bias the reported degradation magnitudes could be overstated.
What would settle it
A prospective blinded study in which board-certified physicians score a large sample of watermarked and unwatermarked outputs on open-ended clinical documentation, not multiple choice, could settle the claim: if expert-rated harm rates do not rise with watermark presence at the strongest detectability settings, the degradation reported here is an artifact of the judge rather than a clinical effect.
If this is right
- Accuracy on medical benchmarks is not a sufficient safety signal: a watermark can raise the faulty-but-correct rate from 11.4% to 26.3% on a 14B model while accuracy drops only 3.2 points.
- Watermarking vision-language models corrupts visual reporting: supported image claims fall and contradicted claims rise in most configurations, even for models whose accuracy stays within the baseline confidence band.
- For reasoning models, watermarking only the final answer is essentially free, while watermarking the reasoning trace inflates output length by up to 69% and degrades terminology and efficiency.
- Degradation is model- and scheme-specific: general-purpose 70B models can be stable while medical-specialised 70Bs collapse, and each scheme damages a different axis; even 'distortion-free' schemes show significant degradation.
- Domain-specific evaluation should be a prerequisite for deploying watermarked models in medicine, because current benchmarks can otherwise obscure clinically consequential failures.
Where Pith is reading between the lines
- If the branching-point mechanism generalizes, watermark harm is concentrated at decision-critical tokens, so a token-importance-aware watermark—preserving clinical content while watermarking low-stakes tokens—could maintain traceability with less clinical damage; the paper leaves this design open.
- The findings imply that regulatory operating points defined only by detectability (for example, 99% true-positive rate at 1% false-positive rate) are insufficient; compliance benchmarks should include domain-quality audits, or a watermark can pass detection requirements while degrading the text clinicians read.
- The judge-validation numbers suggest which axes to trust: fabrication counts and pairwise divergence show substantial agreement with physicians, while misapplied-term and image-contradiction flags show only moderate agreement; a larger prospective physician study on those low-agreement axes would sharpen or revise the reported magnitudes.
- For multi-turn clinical workflows, watermarked outputs become context for subsequent generations; the paper notes it remains unclear whether degradation compounds or attenuates, so testing watermark harm in interactive report-writing and follow-up dialogue settings is a direct extension.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a large-scale empirical evaluation of five text-watermarking schemes (KGW, DipMark, AAR, PPL, SynthID) applied to medical LLMs and VLMs on MedQA and MedXpertQA-MM. The authors introduce three LLM-as-judge protocols—single-response audit, pairwise divergence audit, and reasoning-trace audit—to measure reasoning quality, terminology misuse, hallucinations, and multimodal grounding beyond MCQ accuracy. Their central claim is that watermarking can substantially degrade clinical reasoning even when benchmark accuracy remains within baseline confidence intervals: e.g., faulty-but-correct rates rise from 11.4% to 26.3% on PHI-4-14B under SynthID while accuracy drops only 3.2 pp, and fabricated entities increase by up to 39.2 per 100 questions on LINGSHU-7B. They also report that watermarking reasoning traces in thinking models inflates length and degrades terminology, while restricting the watermark to the final answer preserves reasoning quality. The paper concludes that accuracy-only evaluation is insufficient for safe deployment of watermarked models in medicine.
Significance. If the central claim holds, this is an important result for the trustworthy-deployment literature: it shows that standard accuracy benchmarks can mask clinically consequential watermark-induced degradation, and it identifies answer-only watermarking as a practical mitigation for reasoning models. The study is carefully designed in several respects: paired seeds isolate the watermark effect, bootstrapped CIs are reported, full per-configuration grids are in the appendix, and the judge validation effort is substantial. The branching-point experiment is a useful mechanistic probe. However, the load-bearing audit quantities rest on LLM-judge agreement that is moderate on some headline axes, and the paper does not yet rule out condition-dependent judge bias.
major comments (4)
- [Sec. 4.3, Table 10] The headline deltas in Table 2 and Figure 4 rely on the 'misapplied terms' and 'contradicted image claims' axes. Physician agreement for misapplied terms is only κ=0.42/ρ=0.47, and judge–judge agreement for contradicted image claims is only ρ=0.50. The paper attributes low κ to base-rate sensitivity, but κ is prevalence-dependent and this does not address whether disagreement is condition-dependent. Since the reported quantities are differences between watermarked and unwatermarked outputs, a judge that flags more errors in less-fluent watermarked text would inflate the deltas; paired anchoring removes constant bias but not differential bias. Please report agreement stratified by condition (watermarked vs. unwatermarked), and ideally re-score the headline axes with physician adjudication or a second physician.
- [Sec. 4.3 vs. Introduction/Conclusion] The Introduction and Conclusion state that the judges were validated against 'two board-certified physicians,' but Sec. 4.3 describes a single board-certified physician annotating the 650 samples, and no physician–physician agreement is reported. The acknowledgments thank one physician for annotation. This discrepancy overstates the validation basis. Please clarify exactly how many physicians annotated which subsets, report per-physician agreement if two were used, or remove the 'two physicians' claim.
- [Sec. 4.2, Tables 7 and 8] Statistical significance is reported at p<0.01 for many cells, but the experiment spans 7 models × 5 schemes × roughly 12 axes, yielding hundreds of comparisons. No multiple-testing correction is applied, so some of the scheme-specific 'different axis' conclusions may be false positives. Please report FDR-adjusted q-values or pre-specify a small set of primary axes; at minimum, state the total number of comparisons and discuss the expected number of p<0.01 findings under the null.
- [App. D.2, Figure 12] The clinician validation panel is explicitly not statistically powered (5 scored outputs per configuration), and the text says the deployment operating point is TPR≈99% while the figure caption labels the dotted vertical guide as TPR=80%. This panel is a qualitative check, not a confirmation of the automated judge's dose–response. Please either strengthen the panel (more items, more physicians) or soften the claim that the clinician panel confirms the main findings.
minor comments (4)
- [Abstract vs. Sec. 3.1] The abstract says 11 LLMs, but Sec. 3.1 lists 6 instruction-tuned + 4 reasoning models = 10; BIOMISTRAL-7B appears only in the appendix. Please reconcile the model count.
- [Sec. 4.1, footnote 1] The footnote states BIOMISTRAL-7B collapses with KGW −37.5 and SynthID −14.2 pp, but Table 7 reports −38.7 and −30.4. Please correct the inconsistency.
- [App. D.2, Figure 12 caption] The caption says 'Dotted vertical guide marks TPR = 80%' while the body text discusses TPR≈99% as the deployment point. If intentional, explain; otherwise align the caption with the text.
- [Sec. A.2] The reproducibility section says code is released, but I did not see a repository URL or artefact identifier. Please add the link or a statement about availability.
Circularity Check
No derivative chain to collapse; empirical evaluation with one minor, non-load-bearing self-citation.
full rationale
This is an empirical measurement study, not a derivation-from-principles paper. The central claim (watermarking degrades medical reasoning even when accuracy holds) is established by paired generation experiments across 11 LLMs and 7 VLMs, five watermark schemes, and three LLM-as-judge audits. The judge audits are not defined in terms of watermark status: GEMINI-3-FLASH is blind to condition and gold answer in the single-response audit, and the pairwise audit strips condition labels and randomizes response order; validation is against a human physician (Sec. 4.3, Table 10). No equation in the paper reduces to its own input. The only self-citation is [20] (Gloaguen et al., same research group), which supplies the PPL watermark implementation and an optimality statement for AAR among sum-based distortion-free detectors; both are auxiliary descriptions of the evaluated methods and do not carry the conclusion. Even if [20] were set aside, the empirical contrasts (e.g., PHI-4-14B F&C 11.4%->26.3% under SynthID; LINGSHU-7B +39.2 fabricated entities/100) stand on the reported generation and blind-judge data. The judge-agreement limitations (kappa 0.42 for misapplied terms, rho 0.50 for image-contradicted claims) are a validity or measurement concern, not circularity: the judge is independent of the watermarked models, and the paper transparently reports the weak agreements. Thus no step satisfies the standard of an input-output equivalence by construction; score 1 reflects one minor, non-load-bearing self-citation.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Paired-seed generation isolates the watermark as the only source of divergence between watermarked and unwatermarked completions.
- domain assumption The GEMINI-3-FLASH judge's audit axes, after physician validation on 650 samples, are valid measures of clinical reasoning quality.
- domain assumption TPR ≥99% at 1% FPR is the appropriate regulatory operating point for evaluating watermark degradation.
- domain assumption MedQA and MedXpertQA-MM multiple-choice accuracy is a meaningful proxy for the clinical performance relevant to watermark safety.
- domain assumption The five watermark implementations faithfully match the published schemes.
read the original abstract
Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, underexplored. In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning. Importantly, we complement existing evaluations by introducing a human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations. Our results reveal that watermarking can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings. Notably, we find that the absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to clinical text, can systematically obscure practical watermark-induced degradations. Our findings establish domain-specific evaluation as a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures.
Figures
Reference graph
Works this paper leans on
-
[1]
2026 physician survey on augmented intelligence
American Medical Association. 2026 physician survey on augmented intelligence. https:// www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf , 2026. Accessed: 2026-04-30
2026
-
[2]
Anirudh S. Gowda, Meer C. Chhabria, and Ryan C. Lam. Use of large language models on radiology reports: A scoping review.Journal of the American College of Radiology, 23(3): 437–454, 2026. doi: 10.1016/j.jacr.2025.10.005. Published online 2025
-
[3]
Weissman, Toni Mankowitz, and Genevieve P
Gary E. Weissman, Toni Mankowitz, and Genevieve P. Kanter. Unregulated large language models produce medical device-like output.npj Digital Medicine, 8(1):148, 2025. doi: 10. 1038/s41746-025-01544-y
2025
-
[4]
Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (AI Act)
European Union. Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (AI Act). Official Journal of the European Union, L 2024/1689, 2024
2024
-
[5]
Second draft code of practice on marking and labelling of AI-generated content
European Commission. Second draft code of practice on marking and labelling of AI-generated content. https://digital-strategy.ec.europa.eu/en/library/ commission-publishes-second-draft-code-practice-marking-and-labelling-ai-generated-content ,
-
[6]
WaterBench: Towards holistic evaluation of watermarks for large language models
Shangqing Tu, Yuliang Sun, Yushi Bai, Jifan Yu, Lei Hou, and Juanzi Li. WaterBench: Towards holistic evaluation of watermarks for large language models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, August 2024. Association for Computational Linguistics. URL https://ac...
2024
-
[7]
Johannes Moll, Markus Graf, Tristan Lemke, Nicolas Lenhart, Daniel Truhn, Jean-Benoit Delbrouck, Jiazhen Pan, Daniel Rueckert, Lisa C. Adams, and Keno K. Bressem. Evaluating reasoning faithfulness in medical vision-language models using multimodal perturbations. In Proceedings of Machine Learning for Health (ML4H), volume 297 ofProceedings of Machine Lear...
arXiv 2025
-
[8]
Paul Hager, Friederike Jungmann, Robbie Holland, Kunal Bhagat, Inga Hubrecht, Manuel Knauer, Jakob Vielhauer, Marcus Makowski, Rickmer Braren, Georgios Kaissis, and Daniel Rueckert. Evaluation and mitigation of the limitations of large language models in clinical decision-making.Nature Medicine, 30:2613–2622, 2024. doi: 10.1038/s41591-024-03097-1
-
[9]
What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021
2021
-
[10]
Siun Kim and Hyung-Jin Yoon. Questioning our questions: How well do medical QA bench- marks evaluate clinical capabilities of language models? InProceedings of the 24th Workshop on Biomedical Language Processing (BioNLP), pages 274–296, Vienna, Austria, 2025. Association for Computational Linguistics. doi: 10.18653/v1/2025.bionlp-1.24
-
[11]
Hu et al
Y . Hu et al. Omnimedvqa: A new large-scale comprehensive evaluation benchmark for medical lvlm. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 11
2024
-
[12]
Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, et al. GMAI-MMBench: A comprehensive multimodal evaluation benchmark towards general medical AI.arXiv preprint arXiv:2408.03361, 2024
Pith/arXiv arXiv 2024
-
[13]
MedXpertQA: Benchmarking expert-level medical reasoning and understanding
Yuxin Zuo, Shang Qu, Yifei Li, Zhang-Ren Chen, Xuekai Zhu, Ermo Hua, Kaiyan Zhang, Ning Ding, and Bowen Zhou. MedXpertQA: Benchmarking expert-level medical reasoning and understanding. InProceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 ofProceedings of Machine Learning Research, pages 80961–80990. PMLR, 2025. arXiv:2...
Pith/arXiv arXiv 2025
-
[14]
Umakanta Maharana, Sarthak Verma, Avarna Agarwal, Prakashini Mruthyunjaya, Dwarikanath Mahapatra, Sakir Ahmed, and Murari Mandal. Right prediction, wrong reasoning: Uncovering LLM misalignment in RA disease diagnosis.arXiv preprint arXiv:2504.06581, 2025
Pith/arXiv arXiv 2025
-
[15]
Med-HALT: Medical domain hallucination test for large language models
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Med-HALT: Medical domain hallucination test for large language models. InProceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 314–334, Singapore, 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.conll-1.21
-
[16]
Veysel Kocaman, Mustafa Aytu ˘g Kaya, Andrei Marian Feier, and David Talby. Clinical large language model evaluation by expert review (CLEVER): Framework development and validation.JMIR AI, 4(1):e72153, 2025. doi: 10.2196/72153
-
[17]
Rahul K. Arora, Jason Wei, Rebecca S. Hicks, et al. HealthBench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025
Pith/arXiv arXiv 2025
-
[18]
Aiwei Liu, Leyi Pan, Yijian Lu, Jingjing Li, Xuming Hu, Lijie Wen, Irwin King, and Philip S. Yu. A survey of text watermarking in the era of large language models.ACM Computing Surveys, 57:47:1–47:36, 2024. doi: 10.1145/3691626. URLhttps://arxiv.org/abs/2312.07913
Pith/arXiv arXiv 2024
-
[19]
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. A watermark for large language models. InInternational Conference on Machine Learning, pages 17061–17084. PMLR, 2023
2023
-
[20]
A unified framework for llm watermarks, 2026
Thibaud Gloaguen, Robin Staab, Nikola Jovanovi´c, and Martin Vechev. A unified framework for llm watermarks, 2026. URLhttps://arxiv.org/abs/2602.06754
arXiv 2026
-
[21]
Dipmark: A stealthy, efficient and resilient watermark for large language models
Yihan Wu, Zhengmian Hu, Hongyang Zhang, and Heng Huang. Dipmark: A stealthy, efficient and resilient watermark for large language models. 2023
2023
-
[22]
Watermarking of large language models
Scott Aaronson. Watermarking of large language models. InWorkshop on Large Language Models and Transformers, Simons Institute, UC Berkeley, 2023
2023
-
[23]
Scalable watermarking for identifying large language model outputs.Nature, 634(8035):818–823, 2024
Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovicova, et al. Scalable watermarking for identifying large language model outputs.Nature, 634(8035):818–823, 2024
2024
-
[24]
Robust distortion- free watermarks for language models
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang. Robust distortion- free watermarks for language models. 2024
2024
-
[25]
Peiru Yang, Xintian Li, Wanchun Ni, Jinhua Yin, Huili Wang, Guoshun Nan, Shangguang Wang, Yongfeng Huang, and Tao Qi. Enhancing watermarking quality for LLMs via contextual generation states awareness.arXiv preprint arXiv:2506.07403, 2025
Pith/arXiv arXiv 2025
-
[26]
Who wrote this code? watermarking for code generation
Taehyun Lee, Seokhee Hong, Jaewoo Ahn, Ilgee Hong, Hwaran Lee, Sangdoo Yun, Jamin Shin, and Gunhee Kim. Who wrote this code? watermarking for code generation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, 2024
2024
-
[27]
WatME: Towards lossless watermarking through lexical redundancy
Liang Chen, Yatao Bian, Yang Deng, Deng Cai, Shuaiyi Li, Peilin Zhao, and Kam-Fai Wong. WatME: Towards lossless watermarking through lexical redundancy. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9166–9180, Bangkok, Thailand, August 2024. Association for Computational Linguistic...
-
[28]
Distilling the thought, watermarking the answer: A principle semantic guided watermark for reasoning large language models.OpenReview, 2026
Shuliang Liu, Xingyu Li, Hongyi Liu, Dong Fang, Duan Bingchen, Zheng Qi, Lingfeng Su, and Xuming Hu. Distilling the thought, watermarking the answer: A principle semantic guided watermark for reasoning large language models.OpenReview, 2026. URL https: //openreview.net/forum?id=T6NVogsXCZ. ICLR 2026 Poster
2026
-
[29]
Factuality beyond coherence: Evaluating LLM watermarking methods for medical texts
Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng, and Qian Lou. Factuality beyond coherence: Evaluating LLM watermarking methods for medical texts. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 15129–15147, Suzhou, China, 2025. Association for Computational Linguistics. doi: 10.18653/v1/2025. fin...
arXiv 2025
-
[30]
The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori et al. The Llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Pith/arXiv arXiv 2024
-
[31]
Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025
Gemma Team. Gemma 3 technical report.arXiv preprint arXiv:2503.19786, 2025
Pith/arXiv arXiv 2025
-
[32]
Phi-4 technical report.arXiv preprint arXiv:2412.08905, 2024
Marah Abdin et al. Phi-4 technical report.arXiv preprint arXiv:2412.08905, 2024
Pith/arXiv arXiv 2024
-
[33]
OpenBioLLMs: Advancing open-source large language models for healthcare and life sciences
Ankit Pal and Malaikannan Sankarasubbu. OpenBioLLMs: Advancing open-source large language models for healthcare and life sciences. https://huggingface.co/aaditya/ Llama3-OpenBioLLM-70B, 2024
2024
-
[34]
UltraMedical: Building specialized generalists in biomedicine
Kaiyan Zhang, Sihang Zeng, Ermo Hua, Ning Ding, Zhang-Ren Chen, Zhiyuan Ma, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Jin-Fang Hu, Zhiyuan Liu, and Bowen Zhou. UltraMedical: Building specialized generalists in biomedicine. InAdvances in Neural Infor- mation Processing Systems 37: Datasets and Benchmarks Track (NeurIPS), 2024. Spotlight; arX...
Pith/arXiv arXiv 2024
-
[35]
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. DeepSeek-R1: Incentivizing reason- ing capability in LLMs via reinforcement learning.Nature, 645:633–638, 2025. doi: 10.1038/s41586-025-09422-z
-
[36]
Qwen3.5: Towards native multimodal agents, February 2026
Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URL https: //qwen.ai/blog?id=qwen3.5. Blog post
2026
-
[37]
Qwen3 technical report.arXiv preprint, 2025
Qwen Team. Qwen3 technical report.arXiv preprint, 2025
2025
-
[38]
Gemma 4: Expanding the gemmaverse with Apache 2.0, April 2026
Google Open Source. Gemma 4: Expanding the gemmaverse with Apache 2.0, April 2026. URL https://opensource.googleblog.com/2026/03/ gemma-4-expanding-the-gemmaverse-with-apache-20.html . Google Open Source Blog
2026
-
[39]
Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Cheng- hao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, et al. Lingshu: A generalist foun- dation model for unified multimodal medical understanding and reasoning.arXiv preprint arXiv:2506.07044, 2025
Pith/arXiv arXiv 2025
-
[40]
MedGemma technical report.arXiv preprint arXiv:2507.05201, 2025
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, et al. MedGemma technical report.arXiv preprint arXiv:2507.05201, 2025. 13 A Additional Experimental Details A.1 Watermarking Scheme Details All five schemes share a common structure: at each token posit...
Pith/arXiv arXiv 2025
-
[42]
the answer is X
draws (gu ∈ {0,1}), so that in expectation half of the tokens are green. A constant logit bias δ is added to the green-list tokens before sampling, yielding the soft logit-bias watermark q(g)∝pexp(δ g).(1) This is the soft Kirchenbauer scheme; we do not use the hard variant that excludes red-list tokens entirely. Detection tests whether the proportion of ...
2000
-
[43]
Correctness configuration:BOTH_CORRECT,BOTH_WRONG_SAME_ANSWER,32 BOTH_WRONG_DIFF_ANSWER,R1_ONLY_CORRECT,R2_ONLY_CORRECT, or33 ONE_OR_BOTH_NONE.34
-
[44]
configuration
Diagnostic interpretation divergence(NONE/MINOR/MAJOR): whether the two responses35 make mutually exclusive claims about the same clinical data (e.g., different mechanism attribu-36 tions, image findings, or value interpretations).37 Output schema (abbreviated).38 32 Output JSON - pairwise divergence audit { "configuration": { "response _1_answer", "respo...
-
[2026]
Accessed: 2026-04-30
Published: 2026-03-05. Accessed: 2026-04-30
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.