Pith. sign in

REVIEW 9 cited by

Uncertainty in Natural Language Generation: From Theory to Applications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.15703 v1 pith:BUDC7OON submitted 2023-07-28 cs.CL cs.AIcs.LG

Uncertainty in Natural Language Generation: From Theory to Applications

classification cs.CL cs.AIcs.LG
keywords uncertaintylanguageapplicationsgenerationnaturaltheorysystemsactive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances of powerful Language Models have allowed Natural Language Generation (NLG) to emerge as an important technology that can not only perform traditional tasks like summarisation or translation, but also serve as a natural language interface to a variety of applications. As such, it is crucial that NLG systems are trustworthy and reliable, for example by indicating when they are likely to be wrong; and supporting multiple views, backgrounds and writing styles -- reflecting diverse human sub-populations. In this paper, we argue that a principled treatment of uncertainty can assist in creating systems and evaluation protocols better aligned with these goals. We first present the fundamental theory, frameworks and vocabulary required to represent uncertainty. We then characterise the main sources of uncertainty in NLG from a linguistic perspective, and propose a two-dimensional taxonomy that is more informative and faithful than the popular aleatoric/epistemic dichotomy. Finally, we move from theory to applications and highlight exciting research directions that exploit uncertainty to power decoding, controllable generation, self-assessment, selective answering, active learning and more.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

    cs.RO 2026-05 conditional novelty 8.0

    The paper presents RoboAbstention, a new benchmark showing frontier VLMs and embodied planners abstain on only 16.5-39% of 6,069 instructions grounded in robotics images, with prompting interventions raising rates to ...

  2. Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models

    cs.CL 2026-05 unverdicted novelty 7.0

    SemGrad measures LLM uncertainty via gradients in semantic space using a Semantic Preservation Score to select embeddings, with HybridGrad combining it with parameter gradients to outperform sampling-based baselines e...

  3. Gradients with Respect to Semantics Preserving Embeddings Tell the Uncertainty of Large Language Models

    cs.CL 2026-05 unverdicted novelty 7.0

    SemGrad is a gradient-based uncertainty quantification technique for free-form LLM generation that operates in semantic space using a Semantic Preservation Score to select stable embeddings.

  4. When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

    cs.CL 2026-06 unverdicted novelty 6.0

    Global calibration metrics like ECE are confounded by accuracy; the proposed ACE framework with three accuracy-controlled views shows many prior calibration advantages weaken or reverse.

  5. Uncertainty Decomposition for Clarification Seeking in LLM Agents

    cs.AI 2026-06 unverdicted novelty 6.0

    A prompt-based uncertainty decomposition separates action confidence from request uncertainty to enable clarification seeking in LLM agents, yielding F1 gains of 73% and 36% over baselines on two new underspecified be...

  6. Epistemic Uncertainty Is Not the Reducible Kind

    stat.ML 2026-06 unverdicted novelty 6.0

    The mutual-information measure of epistemic uncertainty is not reducible by additional data, requiring a split into aleatoric, sample-reducible epistemic, and mechanism-reducible epistemic uncertainty.

  7. Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation

    cs.CL 2026-05 unverdicted novelty 6.0

    BAG prompts LLMs to reason over K sampled responses for strategy selection in multi-turn ambiguous QA, improving accuracy and faithfulness to uncertainty over baselines across six models.

  8. Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models

    cs.CL 2025-02 unverdicted novelty 6.0

    Adapts multi-layer token-level Mahalanobis distance with supervised linear regression to yield improved uncertainty scores for LLM truthfulness tasks.

  9. Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

    cs.CL 2024-08 unverdicted novelty 6.0

    A regression model using attention features and recurrent uncertainty scores improves selective generation in LLMs over unsupervised and supervised baselines on ten datasets and three models.