Pith. sign in

REVIEW 1 cited by

Belief Revision: The Adaptability of Large Language Models Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19764 v2 pith:AFTLOFBD submitted 2024-06-28 cs.CL

classification cs.CL
keywords informationmodelsscenariosbeliefbelief-rbeliefsdeltadesigned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The capability to reason from text is crucial for real-world NLP applications. Real-world scenarios often involve incomplete or evolving data. In response, individuals update their beliefs and understandings accordingly. However, most existing evaluations assume that language models (LMs) operate with consistent information. We introduce Belief-R, a new dataset designed to test LMs' belief revision ability when presented with new evidence. Inspired by how humans suppress prior inferences, this task assesses LMs within the newly proposed delta reasoning ($\Delta R$) framework. Belief-R features sequences of premises designed to simulate scenarios where additional information could necessitate prior conclusions drawn by LMs. We evaluate $\sim$30 LMs across diverse prompting strategies and found that LMs generally struggle to appropriately revise their beliefs in response to new information. Further, models adept at updating often underperformed in scenarios without necessary updates, highlighting a critical trade-off. These insights underscore the importance of improving LMs' adaptiveness to changing information, a step toward more reliable AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    AssertBench measures how often LLMs keep the same true/false evaluation of a fact across contradictory user framings, and finds most tested models agree with the user's framing more when they do not know the fact.

Pith tools