The paper formulates LLM-as-judge evaluation as a two-stage missing-data problem and derives sample-size formulas via doubly robust estimators to achieve desired power while allocating more human reviews where LLM predictability is low.
proceedings of the 2008 conference on empirical methods in natural language processing , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
The paper formulates LLM-as-judge evaluation as a two-stage missing-data problem and derives sample-size formulas via doubly robust estimators to achieve desired power while allocating more human reviews where LLM predictability is low.