Pith. sign in

REVIEW 5 cited by

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.17565 v3 pith:QU7G5CQB submitted 2025-04-24 cs.CL

classification cs.CL
keywords trainingreasoningdatamodelsbasedatasetsa-m-teamam-deepseek-distilled-40m
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although large language models (LLMs) have recently achieved remarkable performance on various complex reasoning benchmarks, the academic community still lacks an in-depth understanding of base model training processes and data quality. To address this, we construct a large-scale, difficulty-graded reasoning dataset containing approximately 3.34 million unique queries of varying difficulty levels and about 40 million distilled responses generated by multiple models over several passes. Leveraging pass rate and Coefficient of Variation (CV), we precisely select the most valuable training data to enhance reasoning capability. Notably, we observe a training pattern shift, indicating that reasoning-focused training based on base models requires higher learning rates for effective training. Using this carefully selected data, we significantly improve the reasoning capabilities of the base model, achieving a pass rate of 79.2\% on the AIME2024 mathematical reasoning benchmark. This result surpasses most current distilled models and closely approaches state-of-the-art performance. We provide detailed descriptions of our data processing, difficulty assessment, and training methodology, and have publicly released all datasets and methods to promote rapid progress in open-source long-reasoning LLMs. The dataset is available at: \href{https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M}{https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M}

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Sibling-guided critique and revision of MCTS reasoning traces yields a 30K-sample dataset that matches or beats 590K-sample baselines on the MATH benchmark for 7B models.

  2. JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models

    cs.CL 2025-07 conditional novelty 4.0 of 10

    JT-Math-8B, an open 8B model family trained with a multi-stage math-focused pipeline, reports math benchmark averages above o1-mini and several 7B open models.

  3. MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

    cs.CL 2025-07 conditional novelty 4.0 of 10

    MiroMind-M1 open-sources a two-stage SFT plus RLVR recipe with a new context-aware multi-stage policy optimization (CAMPO) that claims competitive AIME24, AIME25, and MATH500 scores among Qwen-2.5-based models.

  4. X-Intelligence 3.0: Training and Evaluating Reasoning LLM for Semiconductor Display

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A 32B domain-tuned reasoning model beats DeepSeek-R1-671B on proprietary semiconductor display benchmarks, with an LLM-based evaluation framework and domain RAG.

  5. Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.

Pith tools