REVIEW 2 cited by
Stop Measuring Calibration When Humans Disagree
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct. Correctness is commonly estimated against the human majority class. Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies. We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instance-level measures of calibration that capture key statistical properties of human judgements - class frequency, ranking and entropy.
Forward citations
Cited by 2 Pith papers
-
A Framework for Evaluating LLMs Under Task Indeterminacy
When evaluation items admit multiple valid responses, gold-label accuracy underestimates true model performance, and the paper offers bounds on the true performance from partial knowledge.
-
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurio...
Discussion (0). Continue with ORCID to comment.