REVIEW 7 cited by
Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language and Vision-Language Models (LLMs/VLMs) have revolutionized the field of AI by their ability to generate human-like text and understand images, but ensuring their reliability is crucial. This paper aims to evaluate the ability of LLMs (GPT4, GPT-3.5, LLaMA2, and PaLM 2) and VLMs (GPT4V and Gemini Pro Vision) to estimate their verbalized uncertainty via prompting. We propose the new Japanese Uncertain Scenes (JUS) dataset, aimed at testing VLM capabilities via difficult queries and object counting, and the Net Calibration Error (NCE) to measure direction of miscalibration. Results show that both LLMs and VLMs have a high calibration error and are overconfident most of the time, indicating a poor capability for uncertainty estimation. Additionally we develop prompts for regression tasks, and we show that VLMs have poor calibration when producing mean/standard deviation and 95% confidence intervals.
Forward citations
Cited by 7 Pith papers
-
Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation
In two small VLMs, internal token probability detects errors with AUROC up to 0.99 while verbalized confidence stays near 0.9 and performs near chance, except under severe low light where both fail.
-
Visual-Language-Guided Task Planning for Horticultural Robots
A vision-language model drives a simulated greenhouse robot through simple crop-inspection tasks with ~87% success, but long multi-target tasks collapse to under 10% success.
-
Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
In Qwen2-VL under image degradation, scale raises internal error-detection AUROC from 0.80 to 0.98 while verbalized confidence stays weak, and 4-bit hurts the confidence signal far more than accuracy, so a 7B-4bit mod...
-
Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration
C-PEAL trains the active learning selector with a loss that raises entropy for wrong predictions and lowers it for correct ones, improving sample selection for CLIP-style models under prompt learning and LoRA.
-
Uncertainty Quantification for Motor Imagery BCI -- Machine Learning vs. Deep Learning
Classical BCI classifiers (CSP-LDA, MDRM) give better calibrated uncertainty than deep learning, which tends to be overconfident, but deep learning remains more accurate.
-
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
LLM uncertainty quantification should be judged by whether it improves real human decisions, not by calibration scores on trivia benchmarks.
-
Towards Harmonized Uncertainty Estimation for Large Language Models
CUE combines a supervised correctness classifier with existing LLM uncertainty scores to improve indication, balance, and calibration, reporting AUROC and ECE gains across models and datasets.
Discussion (0). Sign in to comment.