Pith. sign in

REVIEW 5 cited by

Perceptions of Linguistic Uncertainty by Language Models and Humans

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15814 v2 pith:HV2QZYPE submitted 2024-07-22 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagemodelsuncertaintyexpressionshumansstatementlinguisticprior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

_Uncertainty expressions_ such as "probably" or "highly unlikely" are pervasive in human language. While prior work has established that there is population-level agreement in terms of how humans quantitatively interpret these expressions, there has been little inquiry into the abilities of language models in the same context. In this paper, we investigate how language models map linguistic expressions of uncertainty to numerical responses. Our approach assesses whether language models can employ theory of mind in this setting: understanding the uncertainty of another agent about a particular statement, independently of the model's own certainty about that statement. We find that 7 out of 10 models are able to map uncertainty expressions to probabilistic responses in a human-like manner. However, we observe systematically different behavior depending on whether a statement is actually true or false. This sensitivity indicates that language models are substantially more susceptible to bias based on their prior knowledge (as compared to humans). These findings raise important questions and have broad implications for human-AI and AI-AI communication.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A large-scale multilingual evaluation of LLM uncertainty estimation methods across 22 languages and 9 models finds that English reasoning closes the UE gap for low-resource languages and that optimal UE method choice ...

  2. Logical forms complement probability in understanding language model (and human) performance

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Logical form (modality, argument type) predicts LLM syllogism accuracy beyond input perplexity, and models show a rejection bias under necessity.

  3. Revisiting Uncertainty Estimation and Calibration of Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Across 80 LLMs on MMLU-Pro, linguistic verbal uncertainty judged by another LLM gives better calibration and error ranking on average than token-probability or numeric self-reported uncertainty, with exceptions.

  4. Position: Uncertainty Quantification Needs Reassessment for Large-language Model Agents

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A position paper arguing that aleatoric/epistemic uncertainty splits fail for LLM agents and proposing underspecification, interaction, and output-based uncertainty research.

  5. The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories

    cs.CL 2025-01 accept novelty 4.0 of 10

    Pretrained language models can serve as credible cognitive science theories only if researchers validate linking hypotheses and avoid pitfalls of commission and omission.

Pith tools