Pith. sign in

REVIEW 2 cited by

Stop Measuring Calibration When Humans Disagree

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.16133 v2 pith:64IGK6YM submitted 2022-10-28 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords calibrationhumanclassmajoritydisagreehumansmeasuringwhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Calibration is a popular framework to evaluate whether a classifier knows when it does not know - i.e., its predictive probabilities are a good indication of how likely a prediction is to be correct. Correctness is commonly estimated against the human majority class. Recently, calibration to human majority has been measured on tasks where humans inherently disagree about which class applies. We show that measuring calibration to human majority given inherent disagreements is theoretically problematic, demonstrate this empirically on the ChaosNLI dataset, and derive several instance-level measures of calibration that capture key statistical properties of human judgements - class frequency, ranking and entropy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Framework for Evaluating LLMs Under Task Indeterminacy

    cs.LG 2024-11 conditional novelty 6.0 of 10

    When evaluation items admit multiple valid responses, gold-label accuracy underestimates true model performance, and the paper offers bounds on the true performance from partial knowledge.

  2. "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurio...

Pith tools