Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.
Title resolution pending
1 Pith paper cite this work, alongside 68 external citations. Polarity classification is still indexing.
1
Pith paper citing it
68
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.