Pith. sign in

REVIEW 2 cited by

Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03439 v3 pith:C5CZKSBZ submitted 2023-04-07 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoninggpt-4logicaldatasetschatgptbenchmarksperformancelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Harnessing logical reasoning ability is a comprehensive natural language understanding endeavor. With the release of Generative Pretrained Transformer 4 (GPT-4), highlighted as "advanced" at reasoning tasks, we are eager to learn the GPT-4 performance on various logical reasoning tasks. This report analyses multiple logical reasoning datasets, with popular benchmarks like LogiQA and ReClor, and newly-released datasets like AR-LSAT. We test the multi-choice reading comprehension and natural language inference tasks with benchmarks requiring logical reasoning. We further construct a logical reasoning out-of-distribution dataset to investigate the robustness of ChatGPT and GPT-4. We also make a performance comparison between ChatGPT and GPT-4. Experiment results show that ChatGPT performs significantly better than the RoBERTa fine-tuning method on most logical reasoning benchmarks. With early access to the GPT-4 API we are able to conduct intense experiments on the GPT-4 model. The results show GPT-4 yields even higher performance on most logical reasoning datasets. Among benchmarks, ChatGPT and GPT-4 do relatively well on well-known datasets like LogiQA and ReClor. However, the performance drops significantly when handling newly released and out-of-distribution datasets. Logical reasoning remains challenging for ChatGPT and GPT-4, especially on out-of-distribution and natural language inference datasets. We release the prompt-style logical reasoning datasets as a benchmark suite and name it LogiEval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  2. Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DZEN, a parallel Dzongkha-English benchmark of 5,161 school science exam questions, shows large LLM accuracy gaps in Dzongkha; adding English translations narrows the gap.

Pith tools