Pith. sign in

REVIEW 12 cited by

To Believe or Not to Believe Your LLM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02543 v2 pith:6UCW3VT3 submitted 2024-06-04 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords uncertaintyepistemiclargeoutputquantificationresponseswhenallows
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore uncertainty quantification in large language models (LLMs), with the goal to identify when uncertainty in responses given a query is large. We simultaneously consider both epistemic and aleatoric uncertainties, where the former comes from the lack of knowledge about the ground truth (such as about facts or the language), and the latter comes from irreducible randomness (such as multiple possible answers). In particular, we derive an information-theoretic metric that allows to reliably detect when only epistemic uncertainty is large, in which case the output of the model is unreliable. This condition can be computed based solely on the output of the model obtained simply by some special iterative prompting based on the previous responses. Such quantification, for instance, allows to detect hallucinations (cases when epistemic uncertainty is high) in both single- and multi-answer responses. This is in contrast to many standard uncertainty quantification strategies (such as thresholding the log-likelihood of a response) where hallucinations in the multi-answer case cannot be detected. We conduct a series of experiments which demonstrate the advantage of our formulation. Further, our investigations shed some light on how the probabilities assigned to a given output by an LLM can be amplified by iterative prompting, which might be of independent interest.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    LLM probability estimates violate the law of total probability across partitions, and subgroup-aggregated estimates often beat direct population-level estimates (the macro fallacy).

  2. SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

    cs.CL 2026-07 conditional novelty 7.0 of 10

    A DETR-style probe distills multi-sample claim uncertainty into single-pass span detection and continuous Mixture-of-Beta scores, outperforming baselines on a new 293K-span benchmark.

  3. Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models

    cs.LG 2026-06 conditional novelty 6.0 of 10

    Hallucination in LLMs is driven by an oracle-invisible “decoding risk” term that grows with scale and causally compounds errors within a response.

  4. Measuring Faithfulness and Abstention: An Automated Pipeline for Evaluating LLM-Generated 3-ply Case-Based Legal Arguments

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An automated LLM-based evaluator finds that eight LLMs rarely hallucinate factors in legal argument generation but often omit relevant factors and usually fail to abstain when no common ground exists.

  5. Calibrating LLMs with Information-Theoretic Evidential Deep Learning

    cs.LG 2025-02 conditional novelty 6.0 of 10

    IB-EDL regularizes evidential deep learning with an information bottleneck to reduce overconfidence in fine-tuned LLMs, improving calibration across multiple benchmarks.

  6. Paradigm-Based Automatic HDL Code Generation Using LLMs

    cs.PL 2025-01 conditional novelty 6.0 of 10

    A paradigm-based workflow with information-list reuse and a two-phase loop improves LLM-generated Verilog pass rates on VerilogEval, with the full-dataset result built from a hybrid of baseline and proposed-method outputs.

  7. Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses

    cs.CL 2024-12 conditional novelty 6.0 of 10

    PROBAR, the estimated probability that a language model's sampled responses are adequate to the prompt, outperforms semantic entropy for selective prediction across ambiguous and open-ended prompts.

  8. Blast Radius

    cs.AI 2026-08 reject novelty 5.0 of 10

    Reversible eviction of dead context, with recurrence-class burial, cuts agentic coding tokens by 17-26% in a synthetic seven-model evaluation.

  9. Hallucination Detection with Small Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    A multi-small-model ensemble with sentence splitting, z-score normalization, and harmonic mean detects hallucinations in RAG answers with a reported 10% F1 gain over single-model baselines.

  10. Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Sampling multiple VLM-generated visual descriptions and letting a text-only LLM vote on the diagnosis improves zero-shot medical image classification on three MedMNIST datasets.

  11. A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A review that organizes LLM uncertainty quantification into token-level, self-verbalized, semantic-similarity, and mechanistic interpretability categories.

  12. Assessing GPT Model Uncertainty in Mathematical OCR Tasks via Entropy Analysis

    cs.IT 2024-12 reject novelty 3.0 of 10

    The paper reports that GPT-4o's token-level uncertainty, computed as the negative log-likelihood of its output, rises monotonically as image resolution falls from 300 to 72 dpi on a single test page.

Pith tools