Pith. sign in

REVIEW 2 cited by

Evaluating the Deductive Competence of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05452 v2 pith:OJUVE6OC submitted 2023-09-11 cs.CL

classification cs.CL
keywords performancellmsreasoninglanguagecontentdeductivefindformat
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The development of highly fluent large language models (LLMs) has prompted increased interest in assessing their reasoning and problem-solving capabilities. We investigate whether several LLMs can solve a classic type of deductive reasoning problem from the cognitive science literature. The tested LLMs have limited abilities to solve these problems in their conventional form. We performed follow up experiments to investigate if changes to the presentation format and content improve model performance. We do find performance differences between conditions; however, they do not improve overall performance. Moreover, we find that performance interacts with presentation format and content in unexpected ways that differ from human performance. Overall, our results suggest that LLMs have unique reasoning biases that are only partially predicted from human reasoning performance and the human-generated language corpora that informs them.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks

    cs.AI 2025-01 conditional novelty 5.0 of 10

    A syllogism-style, five-stage prompting framework improves LLM accuracy on knowledge-based QA over chain-of-thought baselines.

  2. Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models

    cs.CL 2024-12 conditional novelty 3.0 of 10

    A literature review concludes that LLMs show partial humanlike cognitive patterns in bias, reasoning, and creativity, but the evidence is dominated by GPT models and the 'emergence' framing is not directly tested.

Pith tools