Pith. sign in

REVIEW 10 cited by

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15084 v2 pith:D5WNHT7N submitted 2024-12-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords mathmodelsrewardmodelacemathacrossacemath-72b-instructacemath-72b-rm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifying the correct ones. To develop the instruction-tuned math models, we propose a supervised fine-tuning (SFT) process that first achieves competitive performance across general domains, followed by targeted fine-tuning for the math domain using a carefully curated set of prompts and synthetically generated responses. The resulting model, AceMath-72B-Instruct greatly outperforms Qwen2.5-Math-72B-Instruct, GPT-4o and Claude-3.5 Sonnet. To develop math-specialized reward model, we first construct AceMath-RewardBench, a comprehensive and robust benchmark for evaluating math reward models across diverse problems and difficulty levels. After that, we present a systematic approach to build our math reward models. The resulting model, AceMath-72B-RM, consistently outperforms state-of-the-art reward models. Furthermore, when combining AceMath-72B-Instruct with AceMath-72B-RM, we achieve the highest average rm@8 score across the math reasoning benchmarks. We release model weights, training data, and evaluation benchmarks at: https://research.nvidia.com/labs/adlr/acemath

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  2. Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.

  3. AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A two-stage RL recipe, math-only then code-only, substantially lifts math and code reasoning in 7B/14B distilled models and surpasses several distillation-based open models.

  4. EasyMath: A 0-shot Math Benchmark for SLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    EasyMath, a new 0-shot math benchmark for small language models, shows accuracy rising with model size and training, modest chain-of-thought gains, and better consistency at larger scale.

  5. The Majority is not always right: RL training for solution aggregation

    cs.CL 2025-09 conditional novelty 5.0 of 10

    RL-trained aggregation of multiple LLM solutions outperforms majority voting and reward-model selection on AIME and HMMT math benchmarks.

  6. OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A 2.5M-example code reasoning dataset with critique traces enables Qwen2.5-based models to surpass prior open-weight distilled models on LiveCodeBench via test-time self-critique selection.

  7. AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A 7B reasoning model trained with carefully balanced SFT and RL beats prior small models on math and code benchmarks, with the paper documenting scaling and temperature heuristics.

  8. Improving Large Language Models with Concept-Aware Fine-Tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Adding lightweight multi-token auxiliary heads with a weighted future-token loss improves supervised fine-tuning of Llama-3-8B-Instruct across five diverse tasks.

  9. Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem

    cs.CL 2025-06 conditional novelty 5.0 of 10

    One-shot critique fine-tuning, training on critiques of candidate solutions to a single problem, yields large reasoning gains on math and logic benchmarks at far lower compute than one-shot RL.

  10. JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models

    cs.CL 2025-07 conditional novelty 4.0 of 10

    JT-Math-8B, an open 8B model family trained with a multi-stage math-focused pipeline, reports math benchmark averages above o1-mini and several 7B open models.

Pith tools