Pith. sign in

REVIEW 1 cited by

WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05112 v6 pith:R4ES3TLG submitted 2024-09-08 cs.CL

classification cs.CL
keywords detectiontextwatermarkedwaterseekerlargeaccuracydocumentsefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Watermarking algorithms for large language models (LLMs) have attained high accuracy in detecting LLM-generated text. However, existing methods primarily focus on distinguishing fully watermarked text from non-watermarked text, overlooking real-world scenarios where LLMs generate only small sections within large documents. In this scenario, balancing time complexity and detection performance poses significant challenges. This paper presents WaterSeeker, a novel approach to efficiently detect and locate watermarked segments amid extensive natural text. It first applies an efficient anomaly extraction method to preliminarily locate suspicious watermarked regions. Following this, it conducts a local traversal and performs full-text detection for more precise verification. Theoretical analysis and experimental results demonstrate that WaterSeeker achieves a superior balance between detection accuracy and computational efficiency. Moreover, its localization capability lays the foundation for building interpretable AI detection systems. Our code is available at https://github.com/THU-BPM/WaterSeeker.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BiMarker: Enhancing Text Watermark Detection for Large Language Models with Bipolar Watermarks

    cs.LG 2025-01 conditional novelty 6.0 of 10

    BiMarker splits generated text into alternating positive and negative poles and uses the difference in green-token counts to detect LLM watermarks more accurately than KGW.

Pith tools