Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurious and racial biases persist.
Stop Measuring Calibration When Humans Disagree
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct. Correctness is commonly estimated against the human majority class. Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies. We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instance-level measures of calibration that capture key statistical properties of human judgements - class frequency, ranking and entropy.
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations
Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurious and racial biases persist.