REVIEW 16 cited by
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The recent performance leap of Large Language Models (LLMs) opens up new opportunities across numerous industrial applications and domains. However, erroneous generations, such as false predictions, misinformation, and hallucination made by LLMs, have also raised severe concerns for the trustworthiness of LLMs', especially in safety-, security- and reliability-sensitive scenarios, potentially hindering real-world adoptions. While uncertainty estimation has shown its potential for interpreting the prediction risks made by general machine learning (ML) models, little is known about whether and to what extent it can help explore an LLM's capabilities and counteract its undesired behavior. To bridge the gap, in this paper, we initiate an exploratory study on the risk assessment of LLMs from the lens of uncertainty. In particular, we experiment with twelve uncertainty estimation methods and four LLMs on four prominent natural language processing (NLP) tasks to investigate to what extent uncertainty estimation techniques could help characterize the prediction risks of LLMs. Our findings validate the effectiveness of uncertainty estimation for revealing LLMs' uncertain/non-factual predictions. In addition to general NLP tasks, we extensively conduct experiments with four LLMs for code generation on two datasets. We find that uncertainty estimation can potentially uncover buggy programs generated by LLMs. Insights from our study shed light on future design and development for reliable LLMs, facilitating further research toward enhancing the trustworthiness of LLMs.
Forward citations
Cited by 16 Pith papers
-
Hallucination Detection in Large Language Models Using Diversion Decoding
Forcing an LLM away from its greedy answer yields resistance features that train a classifier detecting hallucinations more accurately and cheaply than semantic entropy.
-
Neural Message-Passing on Attention Graphs for Hallucination Detection
CHARM trains graph neural networks on token-attention graphs built from LLM computational traces and outperforms prior hallucination detectors on five benchmarks at token and response level.
-
Can Multiple Responses from an LLM Reveal the Sources of Its Uncertainty?
An auxiliary LLM can diagnose whether an LLM's uncertainty comes from ambiguous questions or missing knowledge by analyzing patterns of disagreement among multiple sampled answers.
-
Shaking to Reveal: Perturbation-Based Detection of LLM Hallucinations
SSP adds a learned, sample-specific noise prompt to an LLM input and scores hallucination by the cosine shift in intermediate representations, outperforming output-confidence baselines on QA benchmarks.
-
Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and Attitudes
LLMs show limited, task-dependent accuracy at choosing correct epistemic modals and attitude verbs in controlled stories, with better performance on necessity and fact statements than on possibility and belief statements.
-
Is Your Model Fairly Certain? Uncertainty-Aware Fairness Evaluation for LLMs
UCerF scores LLM fairness by both correctness and confidence, and SynthBias provides 31,756 gender-occupation coreference samples for benchmark testing.
-
Divide-Then-Align: Honest Alignment based on the Knowledge Boundary of RAG
A post-training method that divides RAG queries into four knowledge quadrants and uses DPO to make models abstain appropriately, improving accuracy and abstention on NQ, TriviaQA, and WebQ.
-
The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling
A multi-sample uncertainty classifier filter improves agent reliability (IDKS) by abstaining on uncertain function calls across open SLMs and BFCL-style benchmarks.
-
SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration
SAFECAST augments hidden-state failure-probe training and conformal calibration with visual and language contrast sets, improving VLA failure detection under distribution shift in several tested settings.
-
TRUST: Test-time Resource Utilization for Superior Trustworthiness
TRUST computes confidence as the angular distance between a test image and a slightly modified, maximally-confident version of it, and claims this ranks predictions monotonically.
-
Improving the Calibration of Confidence Scores in Text Generation Using the Output Distribution's Characteristics
Two probability-only confidence metrics, a top-to-kth beam ratio and a tail-thinness score, improve quality correlation for BART and Flan-T5 on several summarization, translation, and QA datasets.
-
Confidence Estimation for Text-to-SQL in Large Language Models
Consistency-based methods are the most reliable confidence signal for text-to-SQL in black-box LLMs, and executing queries against a database adds a useful correctness signal.
-
Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
Across three VLA models and four simulated manipulation tasks, motion-instability and goal-distance metrics correlate with expert-rated execution quality, showing that binary success rates hide large quality differences.
-
Toward Better Generalisation in Uncertainty Estimators: Leveraging Data-Agnostic Features
Adding data-agnostic probability and entropy features to hidden-state probes improves cross-task generalization in most but not all evaluated transfer pairs.
-
ChemAU: Harness the Reasoning of LLMs in Chemical Research with Adaptive Uncertainty Estimation
ChemAU adds a position penalty to token-level uncertainty estimates so that flagged reasoning steps are corrected by a fine-tuned chemistry model, reporting improved accuracy on GPQA, MMLU-Pro, and SuperGPQA chemistry...
-
HD-NDEs: Neural Differential Equations for Hallucination Detection in LLMs
Modeling the full token-by-token trajectory of LLM hidden states with neural ODEs, CDEs, and SDEs improves hallucination detection by over 14% AUC on a constructed true/false benchmark, though gains shrink on QA datasets.
Discussion (0). Continue with ORCID to comment.