REVIEW 4 cited by
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are displaying emergent abilities for math reasoning tasks,and there is a growing attention on enhancing the ability of open-source LLMs through supervised fine-tuning (SFT).In this paper, we aim to explore a general data strategy for supervised data to help optimize and expand math reasoning ability.Firstly, we determine the ability boundary of reasoning paths augmentation by identifying these paths' minimal optimal set.Secondly, we validate that different abilities of the model can be cumulatively enhanced by Mix of Minimal Optimal Sets of corresponding types of data, while our models MMOS achieve SOTA performance on series base models under much lower construction costs.Besides, we point out GSM-HARD is not really hard and today's LLMs no longer lack numerical robustness.Also, we provide an Auto Problem Generator for robustness testing and educational applications.Our code and data are publicly available at https://github.com/cyzhh/MMOS.
Forward citations
Cited by 4 Pith papers
-
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages
Problem-solving data in continued pretraining improves LLM math reasoning more than general math corpora, and tutorship amplification is the most effective synthesis method.
-
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
An EM-style training loop lets 7B math LLMs learn when to invoke code, improving MATH500 by 11 points and AIME by 9.4 points.
-
Curriculum Demonstration Selection for In-Context Learning
A demonstration selection method that samples one example per difficulty tier improves few-shot language model performance by small and sometimes inconsistent margins.
-
Evaluation of LLMs for mathematical problem solving
A three-model, three-dataset LLM math evaluation using a multi-dimensional reasoning rubric, undermined by contradictory accuracy tables.
Discussion (0). Continue with ORCID to comment.