Pith. sign in

REVIEW 2 cited by

Posterior calibration and exploratory analysis for natural language processing models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1508.05154 v2 pith:CXZEVYUT submitted 2015-08-21 cs.CL

classification cs.CL
keywords analysismodelscalibrationexploratorylanguagenaturalposteriorprocessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many models in natural language processing define probabilistic distributions over linguistic structures. We argue that (1) the quality of a model' s posterior distribution can and should be directly evaluated, as to whether probabilities correspond to empirical frequencies, and (2) NLP uncertainty can be projected not only to pipeline components, but also to exploratory data analysis, telling a user when to trust and not trust the NLP analysis. We present a method to analyze calibration, and apply it to compare the miscalibration of several commonly used models. We also contribute a coreference sampling algorithm that can create confidence intervals for a political event extraction task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rationales Are Not Silver Bullets: Measuring the Impact of Rationales on Model Performance and Reliability

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Rationale-augmented finetuning can hurt accuracy while improving calibration, with the sizes of both effects tied linearly to task difficulty.

  2. Clustered Calibration: Representation-Aware Probability Calibration via Learned Subpopulations

    cs.LG 2025-10 reject novelty 5.0 of 10

    Clustered Calibration groups samples by learned representations and calibrates each cluster separately, with a new cluster-binned ECE claimed to rank models by both calibration and AUC.

Pith tools