Pith. sign in

REVIEW 15 cited by

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12209 v1 pith:ZG7HCOSI submitted 2024-05-20 cs.CL

classification cs.CL
keywords mathbenchllmsmathematicalbenchmarkmathematicsmodelsapplicationcapabilities
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in large language models (LLMs) have showcased significant improvements in mathematics. However, traditional math benchmarks like GSM8k offer a unidimensional perspective, falling short in providing a holistic assessment of the LLMs' math capabilities. To address this gap, we introduce MathBench, a new benchmark that rigorously assesses the mathematical capabilities of large language models. MathBench spans a wide range of mathematical disciplines, offering a detailed evaluation of both theoretical understanding and practical problem-solving skills. The benchmark progresses through five distinct stages, from basic arithmetic to college mathematics, and is structured to evaluate models at various depths of knowledge. Each stage includes theoretical questions and application problems, allowing us to measure a model's mathematical proficiency and its ability to apply concepts in practical scenarios. MathBench aims to enhance the evaluation of LLMs' mathematical abilities, providing a nuanced view of their knowledge understanding levels and problem solving skills in a bilingual context. The project is released at https://github.com/open-compass/MathBench .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning

    cs.CL 2025-01 conditional novelty 7.0 of 10

    Step-aligned in-context learning with a first-try retrieval strategy improves LLM mathematical reasoning over problem-level few-shot prompting on multiple benchmarks.

  2. StatEval: A Comprehensive Benchmark for Large Language Models in Statistics

    cs.CL 2025-10 conditional novelty 6.0 of 10

    StatEval is a new 16,000-question statistics benchmark showing that even strong LLMs score below 60% on research-level statistical proof tasks.

  3. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  4. CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges

    cs.CR 2025-04 conditional novelty 6.0 of 10

    CipherBank evaluates 16 LLMs on 2,358 known-plaintext decryption tasks across 9 ciphers; the best model (Claude-3.5) achieves 45.1% accuracy, and o1 reaches 40.6%.

  5. UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    UGPhysics is a new bilingual benchmark of 5,520 undergraduate physics problems; the strongest tested LLM, OpenAI o1-mini, reaches only 49.8% accuracy.

  6. UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models

    cs.CL 2025-01 conditional novelty 6.0 of 10

    UGMathBench provides a 5,062-problem dynamic benchmark for undergraduate math reasoning and shows that leading LLMs solve all versions of only about half the problems.

  7. TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TReB evaluates 26 large language models on 26 table reasoning subtasks using textual, programmatic, and interleaved reasoning modes, finding that the best model reaches only about 70 on a 0-100 judging scale.

  8. SciDA: Scientific Dynamic Assessor of LLMs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    SciDA is a dynamically initialized, multi-discipline olympiad benchmark that shows LLMs perform substantially worse when problem variables are randomized, which the authors attribute to memorization of fixed numerical...

  9. The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants

    cs.CL 2025-05 reject novelty 5.0 of 10

    Clustering-based routing plus self-consistency voting among ten 7B open models reportedly outranks GPT-4.1 and GPT-4.5 on average over 15 diverse benchmarks.

  10. Computational Experiments in Number Theory

    math.NT 2025-04 reject novelty 5.0 of 10

    LLM reaches >=0.95 accuracy on 60 number theory problems with optimal hints; LightGBM classifier empirically supports Dirichlet conductor conjecture via zero features at 93.9% test accuracy for small q.

  11. UrbanPlanBench: A Comprehensive Urban Planning Benchmark for Evaluating Large Language Models

    cs.CL 2025-04 conditional novelty 5.0 of 10

    Large language models mostly fail a Chinese urban-planning certification benchmark, and a new 30k-instruction dataset yields modest performance gains after fine-tuning.

  12. CoinMath: Harnessing the Power of Coding Instruction for Math LLMs

    cs.CL 2024-12 conditional novelty 5.0 of 10

    CoinMath improves math LLM accuracy by training on GPT-4o-generated code rationales with concise comments, descriptive naming, and hardcoded solutions.

  13. Do LLMs Understand Ambiguity in Text? A Case Study in Open-world Question Answering

    cs.CL 2024-11 conditional novelty 4.0 of 10

    Adding a rephrasing or context-enrichment prompt improves LLM answer similarity on ambiguous QA over naive prompting, though gains are small and not statistically verified.

  14. Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.

  15. Evaluation of LLMs for mathematical problem solving

    cs.AI 2025-05 reject novelty 3.0 of 10

    A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.

Pith tools