Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Chris Callison-Burch; Delip Rao

arxiv: 2606.00093 · v1 · pith:NFBBO4TLnew · submitted 2026-05-25 · 💻 cs.CL · cs.HC· physics.data-an

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

Delip Rao , Chris Callison-Burch This is my paper

classification 💻 cs.CL cs.HCphysics.data-an

keywords agreementcoefficienthandlingbinaryjudgekappareportingabstention

0 comments

read the original abstract

Validating an LLM judge against human annotations usually means reporting several agreement statistics: accuracy, precision, recall, $F_1$, Cohen's $\kappa$, and one or more rank correlations. A survey of 24 recent LLM-as-judge papers finds metric choice entangled with the judgment scale, tie handling, invalid outputs, and abstention handling, and those choices rarely stated. For binary criteria -- the common case in rubric-based evaluation, where each criterion is graded MET or UNMET -- most of the reported numbers are redundant: Pearson's $r$, Spearman's $\rho$, Kendall's $\tau_b$, the phi coefficient $\phi$, and the Matthews Correlation Coefficient all reduce to a single number on non-degenerate binary data, so reporting several of them only creates an illusion of corroborating evidence. Cohen's $\kappa$ is the one agreement coefficient that adds information: it shares $\phi$'s numerator but normalizes differently, and the gap between them measures how far the judge's positive-label rate has drifted from the human's. We then trace what changes when a judge may abstain with a CANNOT_ASSESS verdict: the three common ways of handling abstentions are not interchangeable preprocessing choices but answer different questions, and they break the binary equivalences. The same equivalences reappear, up to a negligible finite-sample correction, for multi-judge ensembles scored with Fleiss' $\kappa$ or Krippendorff's $\alpha$. We close with a reporting checklist that names the judgment scale, the abstention and tie handling mode, coverage, the confusion matrix, and the aggregation level alongside any scalar agreement coefficient.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

How Humans, Bots, and Agents Communicate About Vulnerabilities in Pull Requests
cs.SE 2026-06 unverdicted novelty 2.0

The authors present a registered report outlining their planned large-scale empirical study of vulnerability communication in pull requests by different account types.