Pith. sign in

REVIEW 2 cited by

WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07138 v2 pith:KPHX5IQN submitted 2023-11-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords watermarkinggenerationdetectionevaluatewaterbenchwatermarksevaluationfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

To mitigate the potential misuse of large language models (LLMs), recent research has developed watermarking algorithms, which restrict the generation process to leave an invisible trace for watermark detection. Due to the two-stage nature of the task, most studies evaluate the generation and detection separately, thereby presenting a challenge in unbiased, thorough, and applicable evaluations. In this paper, we introduce WaterBench, the first comprehensive benchmark for LLM watermarks, in which we design three crucial factors: (1) For benchmarking procedure, to ensure an apples-to-apples comparison, we first adjust each watermarking method's hyper-parameter to reach the same watermarking strength, then jointly evaluate their generation and detection performance. (2) For task selection, we diversify the input and output length to form a five-category taxonomy, covering $9$ tasks. (3) For evaluation metric, we adopt the GPT4-Judge for automatically evaluating the decline of instruction-following abilities after watermarking. We evaluate $4$ open-source watermarks on $2$ LLMs under $2$ watermarking strengths and observe the common struggles for current methods on maintaining the generation quality. The code and data are available at https://github.com/THU-KEG/WaterBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking

    cs.CR 2026-08 conditional novelty 6.0 of 10

    WorldMark modulates watermark strength per token using knowledge-graph saliency, improving robust detection and perplexity for MorphMark watermark variants on C4.

  2. BiMarker: Enhancing Text Watermark Detection for Large Language Models with Bipolar Watermarks

    cs.LG 2025-01 conditional novelty 6.0 of 10

    BiMarker splits generated text into alternating positive and negative poles and uses the difference in green-token counts to detect LLM watermarks more accurately than KGW.

Pith tools