Standard, hierarchical, and weighted conformal prediction applied to LLM watermark scores can control false-positive rates when detecting guideline-violating AI edits in simulated classroom essays.
LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an undetectable and elementary watermarking scheme in the closed setting. Also, in the harder open setting, where the adversary has access to most of the model, we propose an unremovable watermarking scheme.
fields
stat.AP 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Watermark in the Classroom: A Conformal Framework for Adaptive AI Usage Detection
Standard, hierarchical, and weighted conformal prediction applied to LLM watermark scores can control false-positive rates when detecting guideline-violating AI edits in simulated classroom essays.