Pith. sign in

REVIEW 6 cited by

Mathematical Capabilities of ChatGPT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.13867 v2 pith:MAULLFVB submitted 2023-01-31 cs.LG cs.AIcs.CL

Mathematical Capabilities of ChatGPT

classification cs.LG cs.AIcs.CL
keywords mathematicalmathematicschatgptdatasetsgpt-4capabilitiesgraduate-levelmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology. In contrast to formal mathematics, where large databases of formal proofs are available (e.g., the Lean Mathematical Library), current datasets of natural-language mathematics, used to benchmark language models, either cover only elementary mathematics or are very small. We address this by publicly releasing two new datasets: GHOSTS and miniGHOSTS. These are the first natural-language datasets curated by working researchers in mathematics that (1) aim to cover graduate-level mathematics, (2) provide a holistic overview of the mathematical capabilities of language models, and (3) distinguish multiple dimensions of mathematical reasoning. These datasets also test whether ChatGPT and GPT-4 can be helpful assistants to professional mathematicians by emulating use cases that arise in the daily professional activities of mathematicians. We benchmark the models on a range of fine-grained performance metrics. For advanced mathematics, this is the most detailed evaluation effort to date. We find that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a mathematical search engine and knowledge base interface. GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty. Contrary to many positive reports in the media about GPT-4 and ChatGPT's exam-solving abilities (a potential case of selection bias), their overall mathematical performance is well below the level of a graduate student. Hence, if your goal is to use ChatGPT to pass a graduate-level math exam, you would be better off copying from your average peer!

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

    cs.CL 2023-07 unverdicted novelty 7.0

    SciBench shows current LLMs reach at most 43.22% accuracy on curated collegiate scientific problems and reveals no prompting strategy dominates across all required skills.

  2. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT

    cs.SE 2023-02 accept novelty 7.0

    The authors present a catalog of prompt patterns that provide reusable solutions to common problems in generating and interacting with outputs from LLMs.

  3. On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

    cs.LG 2026-06 unverdicted novelty 6.0

    Prompt-conditioned LLMs face irreducible error floors from language's limited information capacity and alignment constraints, proven via PAC-Bayes bounds on bilevel cheap-talk games for certain task families.

  4. OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces

    cs.AI 2026-05 unverdicted novelty 6.0

    OPT-BENCH and OPT-Agent evaluate LLM self-optimization in large search spaces, showing stronger models improve via feedback but stay constrained by base capacity and below human performance.

  5. Nothing from Something: Can a Language Model Discover 0?

    cs.AI 2026-06 unverdicted novelty 5.0

    Language models require explicit examples to learn zero in arithmetic but language pretraining halves the examples needed.

  6. The Status Quo and Future of AI-TPACK for Mathematics Teacher Education Students: A Case Study in Chinese Universities

    cs.CY 2025-03 unverdicted novelty 3.0

    Survey of Chinese math teacher trainees finds basic AI-TPACK levels, self-efficacy helps, and strong teaching beliefs may hinder progress.