Pith. sign in

REVIEW 5 cited by

GPT-4 Can't Reason

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.03762 v2 pith:25HBGA6P submitted 2023-07-21 cs.CL

classification cs.CL
keywords gpt-4reasoningproblemsdespiteimprovementperformancereasonability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

GPT-4 was released in March 2023 to wide acclaim, marking a very substantial improvement across the board over GPT-3.5 (OpenAI's previously best model, which had powered the initial release of ChatGPT). However, despite the genuinely impressive improvement, there are good reasons to be highly skeptical of GPT-4's ability to reason. This position paper discusses the nature of reasoning; criticizes the current formulation of reasoning problems in the NLP community, as well as the way in which LLM reasoning performance is currently evaluated; introduces a small collection of 21 diverse reasoning problems; and performs a detailed qualitative evaluation of GPT-4's performance on those problems. Based on this analysis, the paper concludes that, despite its occasional flashes of analytical brilliance, GPT-4 at present is utterly incapable of reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SVAgent: AI Agent for Hardware Security Verification Assertion

    cs.CR 2025-07 conditional novelty 6.0 of 10

    SVAgent is a prompt-engineering framework that decomposes security requirements into sub-questions to generate SystemVerilog assertions with higher reported accuracy and consistency than direct LLM generation.

  2. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  3. Causes in neuron diagrams, and testing causal reasoning in Large Language Models. A glimpse of the future of philosophy?

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Neuron diagrams are used both to test causal reasoning in chatbots and to propose DEF-1, a counterfactual definition of cause claimed to cover more classic cases than previous accounts.

  4. Towards a Comparative Framework for Compositional AI Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A categorical framework for compositional generalisation is applied to DisCoCirc models, showing quantum circuits outperform neural networks on systematicity while neural models overfit more.

  5. Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Standard LLMs make frequent errors on small graph coloring problems, while reasoning models o1-mini and DeepSeek-R1 make fewer but still nonzero errors, and no model reaches perfect accuracy.

Pith tools