Pith. sign in

REVIEW 4 cited by

An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.00799 v1 pith:EUJAYPOX submitted 2024-02-23 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords dataabilityllmsreasoningmathmodelsabilitiesboundary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are displaying emergent abilities for math reasoning tasks,and there is a growing attention on enhancing the ability of open-source LLMs through supervised fine-tuning (SFT).In this paper, we aim to explore a general data strategy for supervised data to help optimize and expand math reasoning ability.Firstly, we determine the ability boundary of reasoning paths augmentation by identifying these paths' minimal optimal set.Secondly, we validate that different abilities of the model can be cumulatively enhanced by Mix of Minimal Optimal Sets of corresponding types of data, while our models MMOS achieve SOTA performance on series base models under much lower construction costs.Besides, we point out GSM-HARD is not really hard and today's LLMs no longer lack numerical robustness.Also, we provide an Auto Problem Generator for robustness testing and educational applications.Our code and data are publicly available at https://github.com/cyzhh/MMOS.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Problem-solving data in continued pretraining improves LLM math reasoning more than general math corpora, and tutorship amplification is the most effective synthesis method.

  2. To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

    cs.AI 2025-02 conditional novelty 5.0 of 10

    An EM-style training loop lets 7B math LLMs learn when to invoke code, improving MATH500 by 11 points and AIME by 9.4 points.

  3. Curriculum Demonstration Selection for In-Context Learning

    cs.CL 2024-11 conditional novelty 5.0 of 10

    A demonstration selection method that samples one example per difficulty tier improves few-shot language model performance by small and sometimes inconsistent margins.

  4. Evaluation of LLMs for mathematical problem solving

    cs.AI 2025-05 reject novelty 3.0 of 10

    A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.

Pith tools