Pith. sign in

REVIEW 3 cited by

Uncertainty in Language Models: Assessment through Rank-Calibration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03163 v2 pith:AS4DAEZV submitted 2024-04-04 cs.CL cs.AIcs.LGstat.ML

classification cs.CLcs.AIcs.LGstat.ML
keywords uncertaintymeasuresconfidencelanguagegenerationhoweverlowermodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Language Models (LMs) have shown promising performance in natural language generation. However, as LMs often generate incorrect or hallucinated responses, it is crucial to correctly quantify their uncertainty in responding to given inputs. In addition to verbalized confidence elicited via prompting, many uncertainty measures ($e.g.$, semantic entropy and affinity-graph-based measures) have been proposed. However, these measures can differ greatly, and it is unclear how to compare them, partly because they take values over different ranges ($e.g.$, $[0,\infty)$ or $[0,1]$). In this work, we address this issue by developing a novel and practical framework, termed $Rank$-$Calibration$, to assess uncertainty and confidence measures for LMs. Our key tenet is that higher uncertainty (or lower confidence) should imply lower generation quality, on average. Rank-calibration quantifies deviations from this ideal relationship in a principled manner, without requiring ad hoc binary thresholding of the correctness score ($e.g.$, ROUGE or METEOR). The broad applicability and the granular interpretability of our methods are demonstrated empirically.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

    stat.ME 2026-07 conditional novelty 7.0 of 10

    A prespecified semantic map plus held-out calibration can turn LLM token probabilities into calibrated posterior estimates over declared states, with bounded error and valid coverage in tested settings.

  2. Reconsidering LLM Uncertainty Estimation Methods in the Wild

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Most LLM uncertainty estimates degrade under distribution shift and adversarial prompts, but simple ensembling of scores at test time improves reliability.

  3. Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    A survey of LLM hallucination research that formalizes hallucination types and argues, via incompleteness and undecidability arguments, that hallucinations cannot be fully eliminated.

Pith tools