Pith. sign in

REVIEW 6 cited by

Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.13439 v2 pith:LC2IURWX submitted 2023-02-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords markerscertaintyepistemicuncertaintyexpressionslanguagemodelaffect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increased deployment of LMs for real-world tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs. Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures. We develop a typology of epistemic markers and inject 50 markers into prompts for question answering. We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%. Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance. Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions. These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models

    cs.CL 2026-07 conditional novelty 8.0 of 10

    Across 45 LLMs, the 'right?' tag effect flips from sycophantic to resistant over four years of releases, while the 'maybe?' tag raises agreement in every model — anti-sycophancy training is grammar-keyed and one-sided.

  2. Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Increasing chain-of-thought reasoning budgets in Llama-3.1-8B produces non-monotonic calibration dynamics, with intermediate reasoning causing the worst overconfidence.

  3. CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

    cs.CL 2023-05 unverdicted novelty 6.0 of 10

    CRITIC improves LLM outputs on question answering, math synthesis, and toxicity reduction by having the model interact with tools to critique and revise its initial generations.

  4. Calibrating Model-Based Evaluation Metrics for Summarization

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    A reference-free proxy scoring framework combined with GIRB calibration produces better-aligned evaluation metrics for summarization and outperforms baselines across seven datasets.

  5. Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    LLM-as-a-Judge systems report confidence that overstates their accuracy, and the paper's TH-Score plus LLM-as-a-Fuser improves calibration.

  6. Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

    cs.AI 2023-08 accept novelty 5.0 of 10

    Survey organizes LLM trustworthiness into seven categories and 29 sub-categories, measures eight sub-categories on popular models, and finds that more aligned models generally score higher but with varying effectiveness.

Pith tools