Pith. sign in

REVIEW 6 cited by

Large Language Models Must Be Taught to Know What They Don't Know

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08391 v3 pith:EFQMBPIH submitted 2024-06-12 cs.LG cs.AIcs.CLstat.ML

classification cs.LGcs.AIcs.CLstat.ML
keywords modelsuncertaintygoodknowlargellmswhenargue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When using large language models (LLMs) in high-stakes applications, we need to know when we can trust their predictions. Some works argue that prompting high-performance LLMs is sufficient to produce calibrated uncertainties, while others introduce sampling methods that can be prohibitively expensive. In this work, we first argue that prompting on its own is insufficient to achieve good calibration and then show that fine-tuning on a small dataset of correct and incorrect answers can create an uncertainty estimate with good generalization and small computational overhead. We show that a thousand graded examples are sufficient to outperform baseline methods and that training through the features of a model is necessary for good performance and tractable for large open-source models when using LoRA. We also investigate the mechanisms that enable reliable LLM uncertainty estimation, finding that many models can be used as general-purpose uncertainty estimators, applicable not just to their own uncertainties but also the uncertainty of other models. Lastly, we show that uncertainty estimates inform human use of LLMs in human-AI collaborative settings through a user study.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning fine-tuning makes LLMs more accurate on answerable problems but worse at abstaining on unanswerable ones, across a new 20-dataset benchmark.

  2. Visual hallucination detection in large vision-language models via evidential conflict

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A feature-level evidential conflict score detects incorrect and hallucinated answers in large vision-language models, and a new PRE-HAL benchmark exposes frequent relation-reasoning failures.

  3. Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Two probability-only confidence metrics, a top-to-kth beam ratio and a tail-thinness score, improve quality correlation for BART and Flan-T5 on several summarization, translation, and QA datasets.

  4. Pretrained LLMs Learn Multiple Types of Uncertainty

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLMs encode multiple dataset-specific linear directions in their hidden states that predict their own answer correctness, and these directions are nearly independent across benchmarks.

  5. Extending Epistemic Uncertainty Beyond Parameters Would Assist in Designing Reliable LLMs

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Bayesian Modeling of Experiments is proposed as a unifying framework for quantifying and reducing the many sources of uncertainty in LLM deployments, beyond abstention.

  6. Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A position paper arguing that aleatoric/epistemic uncertainty splits fail for LLM agents and proposing underspecification, interaction, and output-based uncertainty research.

Pith tools