Matched output limits still produce very different execution-failure mixtures across model families, so benchmark accuracy alone hides how models fail.
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurement-validity problem. Across four judgment datasets, we compare two upgrade paths available in practice: scaling Qwen3 dense judges from 1.7B to 32B parameters and moving across MiniMax M2-M2.7 released APIs. The main pattern is that judge upgrades are not interchangeable: only Qwen3 1.7B to 4B gives a robust adjacent gain, while MiniMax adjacent releases do not. Stronger judges reduce but do not remove position and verbosity bias. Repeated-sample juries add little when errors are correlated. Structured debate can move decisions substantially, but without parser and fallback logs those shifts cannot be attributed to deliberation. We argue that LLM-as-judge reports should include dataset slices, bias probes, error-dependence estimates, and protocol audit trails.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2026 1verdicts
ACCEPT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
Matched output limits still produce very different execution-failure mixtures across model families, so benchmark accuracy alone hides how models fail.