Pith. sign in

REVIEW 7 cited by

Overconfidence is Key: Verbalized Uncertainty Evaluation in Large Language and Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02917 v1 pith:LLEPUGMD submitted 2024-05-05 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords vlmscalibrationllmsuncertaintyabilityerrorlanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language and Vision-Language Models (LLMs/VLMs) have revolutionized the field of AI by their ability to generate human-like text and understand images, but ensuring their reliability is crucial. This paper aims to evaluate the ability of LLMs (GPT4, GPT-3.5, LLaMA2, and PaLM 2) and VLMs (GPT4V and Gemini Pro Vision) to estimate their verbalized uncertainty via prompting. We propose the new Japanese Uncertain Scenes (JUS) dataset, aimed at testing VLM capabilities via difficult queries and object counting, and the Net Calibration Error (NCE) to measure direction of miscalibration. Results show that both LLMs and VLMs have a high calibration error and are overconfident most of the time, indicating a poor capability for uncertainty estimation. Additionally we develop prompts for regression tasks, and we show that VLMs have poor calibration when producing mean/standard deviation and 95% confidence intervals.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    In two small VLMs, internal token probability detects errors with AUROC up to 0.99 while verbalized confidence stays near 0.9 and performs near chance, except under severe low light where both fail.

  2. Visual-Language-Guided Task Planning for Horticultural Robots

    cs.RO 2026-01 conditional novelty 6.0 of 10

    A vision-language model drives a simulated greenhouse robot through simple crop-inspection tasks with ~87% success, but long multi-target tasks collapse to under 10% success.

  3. Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    In Qwen2-VL under image degradation, scale raises internal error-detection AUROC from 0.80 to 0.98 while verbalized confidence stays weak, and 4-bit hurts the confidence signal far more than accuracy, so a 7B-4bit mod...

  4. Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration

    cs.CV 2025-07 conditional novelty 5.0 of 10

    C-PEAL trains the active learning selector with a loss that raises entropy for wrong predictions and lowers it for correct ones, improving sample selection for CLIP-style models under prompt learning and LoRA.

  5. Uncertainty Quantification for Motor Imagery BCI -- Machine Learning vs. Deep Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Classical BCI classifiers (CSP-LDA, MDRM) give better calibrated uncertainty than deep learning, which tends to be overconfident, but deep learning remains more accurate.

  6. From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLM uncertainty quantification should be judged by whether it improves real human decisions, not by calibration scores on trivia benchmarks.

  7. Towards Harmonized Uncertainty Estimation for Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    CUE combines a supervised correctness classifier with existing LLM uncertainty scores to improve indication, balance, and calibration, reporting AUROC and ECE gains across models and datasets.

Pith tools