Pith. sign in

REVIEW 3 cited by

Conditional and Modal Reasoning in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.17169 v4 pith:NOZBMB3W submitted 2024-01-30 cs.CL cs.AIcs.LO

classification cs.CLcs.AIcs.LO
keywords llmsreasoninginferencesbasicconditionalslogicallymakemodals
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The reasoning abilities of large language models (LLMs) are the topic of a growing body of research in AI and cognitive science. In this paper, we probe the extent to which twenty-nine LLMs are able to distinguish logically correct inferences from logically fallacious ones. We focus on inference patterns involving conditionals (e.g., 'If Ann has a queen, then Bob has a jack') and epistemic modals (e.g., 'Ann might have an ace', 'Bob must have a king'). These inferences have been of special interest to logicians, philosophers, and linguists, since they play a central role in the fundamental human ability to reason about distal possibilities. Assessing LLMs on these inferences is thus highly relevant to the question of how much the reasoning abilities of LLMs match those of humans. All the LLMs we tested make some basic mistakes with conditionals or modals, though zero-shot chain-of-thought prompting helps them make fewer mistakes. Even the best performing LLMs make basic errors in modal reasoning, display logically inconsistent judgments across inference patterns involving epistemic modals and conditionals, and give answers about complex conditional inferences that do not match reported human judgments. These results highlight gaps in basic logical reasoning in today's LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Most language models ignore stipulated modal semantics under direct prompting, but reasoning mode can restore sensitivity on a balanced paired benchmark.

  2. Math Natural Language Inference: this should be easy!

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new Math NLI corpus from category theory abstracts plus a ten-model evaluation shows LLM unanimous votes approach human labels (88%) but individual LLMs still make basic math reasoning errors.

  3. Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLMs systematically treat modal words like 'must' as evidence of obligation even in non-obligatory contexts, more strongly than humans, and a few-shot plus reasoning prompt can lower the rate of such judgments.

Pith tools