Pith. sign in

REVIEW 6 cited by

Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06730 v2 pith:5KJX2UVA submitted 2024-01-12 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords expresslanguageresponsesuncertaintiesdownstreamfindinvestigatemodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As natural language becomes the default interface for human-AI interaction, there is a need for LMs to appropriately communicate uncertainties in downstream applications. In this work, we investigate how LMs incorporate confidence in responses via natural language and how downstream users behave in response to LM-articulated uncertainties. We examine publicly deployed models and find that LMs are reluctant to express uncertainties when answering questions even when they produce incorrect responses. LMs can be explicitly prompted to express confidences, but tend to be overconfident, resulting in high error rates (an average of 47%) among confident responses. We test the risks of LM overconfidence by conducting human experiments and show that users rely heavily on LM generations, whether or not they are marked by certainty. Lastly, we investigate the preference-annotated datasets used in post training alignment and find that humans are biased against texts with uncertainty. Our work highlights new safety harms facing human-LM interactions and proposes design recommendations and mitigating strategies moving forward.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. Humans overrely on overconfident language models, across languages

    cs.CL 2025-07 conditional novelty 7.0 of 10

    LLMs produce overconfident-sounding answers in all five tested languages, and bilingual users show the highest overreliance risk in Japanese despite its frequent hedges.

  2. Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

    cs.AI 2025-08 unverdicted novelty 6.0 of 10

    Prospect Theory parameters estimated for LLMs become unstable when epistemic markers are injected into prompts, indicating the model is not robust for decision-making under linguistic uncertainty.

  3. Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies

    cs.HC 2025-02 conditional novelty 5.0 of 10

    Explanations increase user reliance on both correct and incorrect LLM answers, while sources and inconsistent explanations reduce overreliance on incorrect answers in a controlled experiment.

  4. "All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Encoder models trained on noisy classroom ratings look super-human under standard concordance metrics, but generalizability, disattenuation, and hierarchical rater analyses show the apparent advantage is partly spurio...

  5. Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals

    cs.CL 2025-09 conditional novelty 4.0 of 10

    The ratio of agreement to disagreement between a small student model and an LLM correlates with the LLM's annotation accuracy across ten datasets and can heuristically select better models.

  6. Bridging Expertise Gaps: The Role of LLMs in Human-AI Collaboration for Cybersecurity

    cs.CR 2025-05 conditional novelty 4.0 of 10

    In n=58 non-expert participants, human-AI collaboration improved phishing precision and intrusion recall, with confident LLM responses strongly influencing user decisions.

Pith tools