The paper formulates LLM-as-judge evaluation as a two-stage missing-data problem and derives sample-size formulas via doubly robust estimators to achieve desired power while allocating more human reviews where LLM predictability is low.
proceedings of the 2008 conference on empirical methods in natural language processing , pages=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
On five tabular security datasets at 10% labels, tuning only the classifier with Bayesian optimization recovers a median 86% of the gains from full joint SSL-classifier optimization.
citing papers explorer
-
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
The paper formulates LLM-as-judge evaluation as a two-stage missing-data problem and derives sample-size formulas via doubly robust estimators to achieve desired power while allocating more human reviews where LLM predictability is low.
-
SemiScope: Disentangling Classifier Tuning and Joint Optimization in Semi-Supervised Security Classification
On five tabular security datasets at 10% labels, tuning only the classifier with Bayesian optimization recovers a median 86% of the gains from full joint SSL-classifier optimization.