REVIEW 3 cited by
Conditional and Modal Reasoning in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The reasoning abilities of large language models (LLMs) are the topic of a growing body of research in AI and cognitive science. In this paper, we probe the extent to which twenty-nine LLMs are able to distinguish logically correct inferences from logically fallacious ones. We focus on inference patterns involving conditionals (e.g., 'If Ann has a queen, then Bob has a jack') and epistemic modals (e.g., 'Ann might have an ace', 'Bob must have a king'). These inferences have been of special interest to logicians, philosophers, and linguists, since they play a central role in the fundamental human ability to reason about distal possibilities. Assessing LLMs on these inferences is thus highly relevant to the question of how much the reasoning abilities of LLMs match those of humans. All the LLMs we tested make some basic mistakes with conditionals or modals, though zero-shot chain-of-thought prompting helps them make fewer mistakes. Even the best performing LLMs make basic errors in modal reasoning, display logically inconsistent judgments across inference patterns involving epistemic modals and conditionals, and give answers about complex conditional inferences that do not match reported human judgments. These results highlight gaps in basic logical reasoning in today's LLMs.
Forward citations
Cited by 3 Pith papers
-
Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?
Most language models ignore stipulated modal semantics under direct prompting, but reasoning mode can restore sensitivity on a balanced paired benchmark.
-
Math Natural Language Inference: this should be easy!
A new Math NLI corpus from category theory abstracts plus a ten-model evaluation shows LLM unanimous votes approach human labels (88%) but individual LLMs still make basic math reasoning errors.
-
Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
LLMs systematically treat modal words like 'must' as evidence of obligation even in non-obligatory contexts, more strongly than humans, and a few-shot plus reasoning prompt can lower the rate of such judgments.
Discussion (0). Continue with ORCID to comment.