Best-of-K selection against an ensemble of AI judges overstates quality at most like the square root of log K times common-mode error, which disagreement-based audits cannot see.
International Conference on Learning Representations , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees
Best-of-K selection against an ensemble of AI judges overstates quality at most like the square root of log K times common-mode error, which disagreement-based audits cannot see.