REVIEW 2 cited by
Evaluating the Deductive Competence of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The development of highly fluent large language models (LLMs) has prompted increased interest in assessing their reasoning and problem-solving capabilities. We investigate whether several LLMs can solve a classic type of deductive reasoning problem from the cognitive science literature. The tested LLMs have limited abilities to solve these problems in their conventional form. We performed follow up experiments to investigate if changes to the presentation format and content improve model performance. We do find performance differences between conditions; however, they do not improve overall performance. Moreover, we find that performance interacts with presentation format and content in unexpected ways that differ from human performance. Overall, our results suggest that LLMs have unique reasoning biases that are only partially predicted from human reasoning performance and the human-generated language corpora that informs them.
Forward citations
Cited by 2 Pith papers
-
SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks
A syllogism-style, five-stage prompting framework improves LLM accuracy on knowledge-based QA over chain-of-thought baselines.
-
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models
A literature review concludes that LLMs show partial humanlike cognitive patterns in bias, reasoning, and creativity, but the evidence is dominated by GPT models and the 'emergence' framing is not directly tested.
Discussion (0). Continue with ORCID to comment.