Pith. sign in

REVIEW 5 cited by

On Subjective Uncertainty Quantification and Calibration in Natural Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05213 v2 pith:AN4IXFBZ submitted 2024-06-07 cs.CL cs.AIcs.LGstat.ML

On Subjective Uncertainty Quantification and Calibration in Natural Language Generation

classification cs.CL cs.AIcs.LGstat.ML
keywords uncertaintycalibrationlanguagequantificationassumptiondataepistemicgeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Applications of large language models often involve the generation of free-form responses, in which case uncertainty quantification becomes challenging. This is due to the need to identify task-specific uncertainties (e.g., about the semantics) which appears difficult to define in general cases. This work addresses these challenges from a perspective of Bayesian decision theory, starting from the assumption that our utility is characterized by a similarity measure that compares a generated response with a hypothetical true response. We discuss how this assumption enables principled quantification of the model's subjective uncertainty and its calibration. We further derive a measure for epistemic uncertainty, based on a missing data perspective and its characterization as an excess risk. The proposed methods can be applied to black-box language models. We illustrate the methods on question answering and machine translation tasks. Our experiments provide a principled evaluation of task-specific calibration, and demonstrate that epistemic uncertainty offers a promising deferral strategy for efficient data acquisition in in-context learning.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Concentration and Calibration in Predictive Bayesian Inference

    stat.ME 2026-05 unverdicted novelty 6.0

    Predictive Bayesian inference posteriors concentrate onto a forward-model-dependent quantity and produce miscalibrated credible sets unless the predictive model contains the true data-generating process.

  2. Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems

    cs.LG 2025-06 unverdicted novelty 6.0

    Introduces a Bayesian framework viewing LLM prompts as textual parameters and proposes MHLP, a novel MCMC algorithm using LLM proposals, to perform inference and improve accuracy plus uncertainty quantification on benchmarks.

  3. Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation

    cs.CL 2026-06 unverdicted novelty 5.0

    Empirical study finds verbalized per-token confidence methods in LLMs for MT perform similarly to internal signals on error detection and calibration but show little correlation.

  4. Trustworthy deep domain adaptation for wearable photoplethysmography signal analysis with decision-theoretic uncertainty quantification

    cs.LG 2026-04 unverdicted novelty 5.0

    Decision-theoretic uncertainty quantification formalizes evaluation of generative domain adaptation trustworthiness for PPG-based atrial fibrillation classification by linking uncertainty to downstream task utility.

  5. Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning

    cs.CL 2026-04 unverdicted novelty 5.0

    Supervised fine-tuning degrades the correlation between confidence scores and output quality in language models, driven by factors like training distribution similarity rather than true quality.