Pith. sign in

REVIEW 4 cited by

GPT-4 Can't Reason

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.03762 v2 pith:25HBGA6P submitted 2023-07-21 cs.CL

classification cs.CL
keywords gpt-4reasoningproblemsdespiteimprovementperformancereasonability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

GPT-4 was released in March 2023 to wide acclaim, marking a very substantial improvement across the board over GPT-3.5 (OpenAI's previously best model, which had powered the initial release of ChatGPT). However, despite the genuinely impressive improvement, there are good reasons to be highly skeptical of GPT-4's ability to reason. This position paper discusses the nature of reasoning; criticizes the current formulation of reasoning problems in the NLP community, as well as the way in which LLM reasoning performance is currently evaluated; introduces a small collection of 21 diverse reasoning problems; and performs a detailed qualitative evaluation of GPT-4's performance on those problems. Based on this analysis, the paper concludes that, despite its occasional flashes of analytical brilliance, GPT-4 at present is utterly incapable of reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SVAgent: AI Agent for Hardware Security Verification Assertion

    cs.CR 2025-07 conditional novelty 6.0 of 10

    SVAgent is a prompt-engineering framework that decomposes security requirements into sub-questions to generate SystemVerilog assertions with higher reported accuracy and consistency than direct LLM generation.

  2. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  3. Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Neuron diagrams are used both to test causal reasoning in chatbots and to propose DEF-1, a counterfactual definition of cause claimed to cover more classic cases than previous accounts.

  4. Towards a Comparative Framework for Compositional AI Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A categorical framework for compositional generalisation is applied to DisCoCirc models, showing quantum circuits outperform neural networks on systematicity while neural models overfit more.

Pith tools