REVIEW 4 major objections 8 minor 21 cited by
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
T0 review · 4 major / 8 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A 7-billion-parameter model trained with reasoning-dense data and graded code rewards beats o1-mini on math and code benchmarks.
desk verdict Real 7B reasoning recipe with public checkpoints and two genuinely new RL ideas, but the o1-mini headline depends on unverifiable decontamination and the paper overclaims general reasoning. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by four mechanisms. The three-stage data mixture builds a general curated distribution, then raises math and code to about 70 percent of tokens, then adds about 10 percent synthetic reasoning responses while extending the context window to 32,768 tokens. Multi-token prediction (MTP) is an auxiliary objective that predicts several future tokens at once and, at inference, is reused for speculative decoding, with the first MTP layer reported to accept about 90 percent of drafted tokens on AIME 2024. The test-difficulty-driven code reward, inspired by International Olympiad in Informatics scoring, clusters each problem's test cases into difficulty levels by their pass rates over many model rollouts and returns the sum of per-group scores, so a solution earns partial credit for passing harder subtasks instead of an all-or-nothing verdict. Easy-data re-sampling keeps perfectly solved problems in a pool that is sampled with probability 10 percent, which stabilizes policy updates in late RL training. These sit on top of GRPO modifications the paper adopts from recent work: removal of the KL loss, dynamic sampling that filters prompts with pass rates of 0 or 1, and an asymmetric clip bound.
What would settle it
Run MiMo-7B-RL on a competition set released after its training cutoff — the next AIME or a LiveCodeBench window that postdates the model — and compare pass@32 with o1-mini; if the reported lead shrinks or inverts on the fresh set while the paper's benchmarks stay high, the advantage is contamination rather than reasoning ability. A cheaper inspection is to search the released RL and SFT data for near-duplicates of evaluation problems at an 8-gram threshold and read the retrieved passages for paraphrased restatements.
Extended reading notes
Core claim
MiMo-7B-RL, a 7B model trained from scratch and then tuned on verifiable problems, reaches 55.4 on AIME 2025, 57.8 on LiveCodeBench v5, and 49.3 on LiveCodeBench v6, surpassing o1-mini (50.7, 53.8, 46.8) on all three while remaining below it on MMLU-Pro and IFEval. The authors attribute the outcome to two linked investments. Pretraining targets what they call reasoning pattern density: extraction tooling that preserves equations and code, small fine-tuned taggers that replace heuristic filters, synthetic reasoning data, and a three-stage mixture whose middle stage devotes about 70 percent of tokens to math and code, with a final stage that adds roughly 10 percent synthetic reasoning responses and extends the context window. Post-training applies a GRPO-style reinforcement learning recipe with rule-based rewards only, no KL term, dynamic sampling, and a test-difficulty-driven code reward that clusters test cases by pass rate and awards partial credit so hard problems yield learnable signal. They also report that RL applied directly to the base model outperforms RL trained on a 32B base model on both math and code, which they take as evidence that reasoning potential is largely set during pretraining.
Load-bearing premise
The headline comparisons assume that AIME 2025 and LiveCodeBench v5 and v6 problems did not appear, in original or paraphrased form, in the 25-trillion-token pretraining corpus, the multi-million-sample SFT set, or the 130K RL problem set, since the only stated defense is a 16-gram overlap filter that catches verbatim but not disguised copies.
Editorial extensions
If this is right
- A 7B open model can exceed a proprietary reasoning model on math and code benchmarks using only rule-based rewards, with no format or length penalties.
- The ceiling of RL training is set largely by the base model: RL run directly on MiMo-7B-Base outperforms RL run on a 32B base on both math and code, suggesting curated pretraining can substitute for scale.
- MTP pays for itself twice: it contributes a training signal and, via speculative decoding, cuts the latency of the long generations that reasoning models produce.
- Scaling SFT from 500K to 6M instances improves reasoning and dialogue and does not blunt later RL gains; a follow-on checkpoint trained with on-policy RL and a longer generation budget reaches 80.1 on AIME 2024.
- The o1-mini advantage is specific to verifiable math and code reasoning: on MMLU-Pro and IFEval, MiMo-7B-RL scores below o1-mini, so the headline comparison is not a claim of general superiority.
Reading between the lines
- Because the paper reports no ablation of the three-stage mixture, a controlled study that varies only the stage-2 math/code ratio or the stage-3 synthetic share would say which stage actually drives the base model's pass@k advantage.
- The pass-rate clustering behind the code reward should transfer to any domain with verifiable sub-results — theorem proving with per-tactic checks, agent tasks with sub-goal verification, scientific computation with partial output checks — wherever all-or-nothing rewards are too sparse.
- The reported margins over o1-mini may not transfer to future contest sets: the 16-gram decontamination filter removes verbatim overlaps only, so disguised or paraphrased evaluation problems inside the 25T-token corpus would inflate the paper's numbers while leaving freshly written competitions as the true test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a full-stack recipe for building MiMo-7B, a 7B-parameter reasoning-focused LLM trained from scratch on 25 trillion tokens. The pretraining contribution is a three-stage data mixture with synthetic reasoning data, an MTP auxiliary objective, and a data-processing pipeline aimed at increasing reasoning-pattern density. The posttraining contribution is an SFT stage followed by GRPO-based RL on 130K verifiable math and code problems, with a test-difficulty-driven code reward, dynamic sampling, and an engineered rollout engine. The paper claims that MiMo-7B-Base outperforms comparable open 7B/8B/9B base models on several reasoning benchmarks, that MiMo-7B-RL-Zero surpasses a 32B RL baseline, and that MiMo-7B-RL beats OpenAI o1-mini on AIME 2025 and LiveCodeBench v5/v6 while maintaining competitive general performance. The open-sourced checkpoints and detailed hyperparameters are notable assets, but the headline comparison depends on decontamination assurances and an internally evaluated benchmark protocol that are not fully documented.
Significance. If the claims survive scrutiny, the paper is significant: a 7B model trained with an openly described recipe would outperform o1-mini on selected math and code benchmarks, at a scale where such results are still uncommon, and would provide a reproducible open-source reference for both pretraining and RL posttraining. The MTP-based speculative decoding with reported acceptance rates is a practical inference contribution, and the Seamless Rollout Engine addresses a real systems bottleneck in RL training. Credit is also due for releasing model checkpoints and for reporting repeated-sampling averages on several benchmarks rather than single draws. However, the general-reasoning claim in the abstract is contradicted by Table 4, the head-to-head comparison against o1-mini depends on unverified decontamination of private training data, and the evaluation protocol for the headline numbers needs clarification. These issues do not necessarily invalidate the core empirical findings, but they must be corrected before the central claims can be accepted.
major comments (4)
- [Section 2.1, Section 3.1, Section 3.2, Tables 4-5] This is a correctness risk rather than a demonstration of contamination, but it is the weakest load-bearing link in the headline comparison.
- [Abstract, Section 1, Table 4] This is not merely a wording issue: the current text makes a general-reasoning claim that the paper's own table refutes.
- [Section 3.5.1, Table 4, Figure 1] This directly affects the validity of the o1-mini margin reported in the abstract.
- [Section 3.3.1, Figure 5] Without this information, the reader cannot judge whether the reported gains on LiveCodeBench come from the reward scheme or from other components of the RL pipeline.
minor comments (8)
- [Section 3.5.2]
- [Section 3.2]
- [Section 3.1]
- [Section 2.2]
- [References]
- [Table 6]
- [Table 2 and Table 3]
- [Figure 5]
Circularity Check
No significant circularity: the report's central claims are empirical measurements on external benchmarks, not derived predictions.
full rationale
The paper makes no derived 'prediction' that reduces to a fitted constant or to a self-citation chain. Pre-training, SFT, and RL are reported as empirical recipes, and the headline comparisons in Table 4 are measured scores on external benchmarks (AIME 2024/2025, LiveCodeBench v5/v6, MMLU-Pro, IFEval) against o1-mini and other models. The test-difficulty reward is calibrated from rollout pass rates, but that calibrates a training signal rather than an evaluation claim, and the reported gains are evaluated on held-out benchmarks with an internal n-gram decontamination step. The only same-team citation (MiMo-VL, Section 3.6) is used to motivate an on-policy RL variant and is not load-bearing: the paper supplies its own training curves and ablations in support. Data contamination would be a correctness risk, not a circularity, because the decontamination is internal and not independently audited; no equation in the paper is equivalent to its inputs by construction.
Assumptions & free parameters
free parameters (6)
- Stage 2 math/code data fraction =
~70%
- Stage 3 synthetic response fraction =
~10%
- MTP loss weight schedule =
0.3 for first 10.3T tokens, then 0.1
- Easy-data resampling probability =
10% (alpha=0.1)
- Easy problem filtering passrate threshold =
90% pass rate over 16 rollouts
- Test-difficulty grouping thresholds and score weights =
not reported
assumptions (5)
- domain assumption Pass@k curves measure a base model's 'reasoning potential' and predict its ceiling under RL.
- domain assumption Verifiable math and code problems with rule-based rewards are a sufficient and robust training signal.
- domain assumption The evaluation benchmarks are not contaminated by the pretraining corpus, SFT data, or RL problem set.
- ad hoc to paper Synthetic reasoning data can be trained for extremely high epochs without overfitting.
- ad hoc to paper Test-case pass rates computed from a few models give stable difficulty groupings for reward shaping.
Cite this review
Pith. "Pith review of MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining." pith.science (2026). https://pith.science/paper/H37IGTXU
@misc{pith2026250507608,
author = {Pith},
title = {Pith review of: MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining},
year = {2026},
howpublished = {\url{https://pith.science/paper/H37IGTXU}},
note = {Machine review of arXiv:2505.07608}
}
read the original abstract
We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing strategy to strengthen the base model's reasoning potential. MiMo-7B-Base is pre-trained on 25 trillion tokens, with additional Multi-Token Prediction objective for enhanced performance and accelerated inference speed. During post-training, we curate a dataset of 130K verifiable mathematics and programming problems for reinforcement learning, integrating a test-difficulty-driven code-reward scheme to alleviate sparse-reward issues and employing strategic data resampling to stabilize training. Extensive evaluations show that MiMo-7B-Base possesses exceptional reasoning potential, outperforming even much larger 32B models. The final RL-tuned model, MiMo-7B-RL, achieves superior performance on mathematics, code and general reasoning tasks, surpassing the performance of OpenAI o1-mini. The model checkpoints are available at https://github.com/xiaomimimo/MiMo.
Forward citations
Cited by 21 Pith papers
-
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction
AdaMTP uses entropy-based segmentation to adaptively mask multi-token prediction losses, improving quality and speed over fixed-horizon multi-token prediction.
-
rStar2-Agent: Agentic Reasoning Technical Report
A 14B model trained with agentic RL and a resample-on-correct rollout strategy scores 80.6% on AIME24 and 69.8% on AIME25, nearly matching DeepSeek-R1 (671B) in one week on 64 GPUs.
-
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
LLMs display stable Participation and Proactiveness risk profiles in multi-agent poker that remain largely robust to opponent mix and adapt heterogeneously under global and personal risk pressure.
-
OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis
OralAgent, a ReAct-style dental agent with 22 vision tools and a 134.8M-token textbook RAG corpus, reaches SOTA on MMOral-Uni, MMOral-OPG, and OralQA-ZH.
-
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
Perception, not reasoning, is the main bottleneck for MLLM STEM visual reasoning, and training on executable reconstruction code measurably fixes it.
-
PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention
Sharing one proxy full-prefix scan across nearby query groups matches dense DSA indexer accuracy while accelerating indexing up to 4× and end-to-end latency up to 1.6×.
-
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
RefBench-PRO organizes REC into attribute, position, interaction, relation, commonsense, and reject tasks; no tested MLLM exceeds 72%, and Ref-R1 raises Qwen2.5-VL-7B from 57.6 to 69.4 on it.
-
Efficiency of turbulence
The efficiency of turbulence, the fraction of input energy stored in the flow, appears bounded and may saturate in a power-law manner across several turbulent flows.
-
Are Reasoning Models More Prone to Hallucination?
Post-training pipeline choice (SFT+RL vs RL-only vs SFT-only) reliably shifts hallucination rates in large reasoning models on fact-seeking benchmarks.
-
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
X3-OPD improves audio-grounded reasoning by training the audio student on its own rollouts with token-level teacher feedback, using a three-tier paired text-audio corpus.
-
Self-Improving is Often Sudden: Enlightenment-style Finetuning for Large-Scale Models
A training-free intervention that scales VLM residual connections and mixes LLM attention heads claims 1–3 point zero-shot accuracy gains, but the evidence is partly tuned on the reported benchmarks.
-
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
Scene-aware multi-agent document synthesis plus error-driven hard-example expansion improves compact Qwen3-VL models on constrained and open-category KIE, topping reported on-device baselines.
-
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
A new benchmark claims to be the first to test VLMs on both external and in-cabin driving risks, and reports a fine-tuned model far outperforming all baselines.
-
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
A 7B reasoning model trained with carefully balanced SFT and RL beats prior small models on math and code benchmarks, with the paper documenting scaling and temperature heuristics.
-
MiniCPM4: Ultra-Efficient LLMs on End Devices
MiniCPM4-8B reportedly matches Qwen3-8B on standard benchmarks while using about 22% of the training tokens, and achieves large long-context speedups on edge devices.
-
MiMo-VL Technical Report
MiMo-VL-7B-RL, a 7B open-source vision-language model, reports state-of-the-art results on 35 of 40 benchmarks and a 59.4 OlympiadBench score, with the report crediting long-CoT pretraining data and mixed on-policy RL.
-
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
One-shot critique fine-tuning, training on critiques of candidate solutions to a single problem, yields large reasoning gains on math and logic benchmarks at far lower compute than one-shot RL.
-
General-Reasoner: Advancing LLM Reasoning Across All Domains
Zero-RL training with a compact generative answer verifier and diverse verifiable questions improves LLM reasoning on general-domain benchmarks while retaining mathematical ability.
-
Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency
Suppressing "Wait"-like reflection tokens at decode time reduces reasoning token counts by 27-51% across five R1-style model families, with mixed accuracy effects.
-
Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
A comprehensive review that categorizes methods for shortening and adaptively triggering chain-of-thought reasoning in large language models.
-
Reinforcement Learning from Human Feedback
The book introduces the origins, mathematical setup, and optimization stages of RLHF including reward modeling, reinforcement learning, rejection sampling, and direct alignment algorithms.
Reference graph
Works this paper leans on
-
[1]
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai. GQA : Training generalized multi-query transformer models from multi-head checkpoints. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4895--4901, Singapore, 2023. Association for Comp...
-
[2]
Claude 3.7 sonnet and claude code, 2025
Anthropic. Claude 3.7 sonnet and claude code, 2025. URL https://www.anthropic.com/claude/sonnet
work page 2025
- [3]
-
[4]
A. Barbaresi. Trafilatura: A web scraping library and command-line tool for text discovery and extraction. In H. Ji, J. C. Park, and R. Xia, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations, pages 122--131, Onli...
-
[5]
Y. Bisk, R. Zellers, R. LeBras, J. Gao, and Y. Choi. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intellige...
work page 2020
-
[6]
A. Z. Broder. On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pages 21--29. IEEE, 1997
work page 1997
-
[7]
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. Evaluating large language models trained on code. ArXiv preprint, abs/2107.03374, 2021. URL https://arxiv.org/abs/2107.03374
arXiv 2021
- [8]
Show all 70 references
-
[9]
Cobbe, V
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. ArXiv preprint, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168
2021 arXiv
-
[10]
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier. Language modeling with gated convolutional networks. In D. Precup and Y. W. Teh, editors, Proceedings of the 34th International Conference on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 , volume 70 of P...
2017
-
[11]
X. Du, Y. Yao, K. Ma, B. Wang, T. Zheng, K. Zhu, M. Liu, Y. Liang, X. Jin, Z. Wei, et al. Supergpqa: Scaling llm evaluation across 285 graduate disciplines. ArXiv preprint, abs/2502.14739, 2025. URL https://arxiv.org/abs/2502.14739
2025 arXiv
-
[12]
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner. DROP : A reading comprehension benchmark requiring discrete reasoning over paragraphs. In J. Burstein, C. Doran, and T. Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of th...
2019 doi
-
[13]
A. P. Gema, J. O. J. Leang, G. Hong, A. Devoto, A. C. M. Mancino, R. Saxena, X. He, Y. Zhao, X. Du, M. R. G. Madani, et al. Are we done with mmlu? ArXiv preprint, abs/2406.04127, 2024. URL https://arxiv.org/abs/2406.04127
2024 arXiv
-
[14]
Gloeckle, B
F. Gloeckle, B. Y. Idrissi, B. Rozi \` e re, D. Lopez - Paz, and G. Synnaeve. Better & faster large language models via multi-token prediction. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL...
2024
-
[15]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. ArXiv preprint, abs/2407.21783, 2024. URL https://arxiv.org/abs/2407.21783
2024 arXiv
-
[16]
A. Gu, B. Rozi \` e re, H. J. Leather, A. Solar - Lezama, G. Synnaeve, and S. Wang. Cruxeval: A benchmark for code reasoning, understanding and execution. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net...
2024
-
[17]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. ArXiv preprint, abs/2501.12948, 2025. URL https://arxiv.org/abs/2501.12948
2025 arXiv
-
[18]
J. He, J. Liu, C. Y. Liu, R. Yan, C. Wang, P. Cheng, X. Zhang, F. Zhang, J. Xu, W. Shen, S. Li, L. Zeng, T. Wei, C. Cheng, B. An, Y. Liu, and Y. Zhou. Skywork open reasoner series. https://capricious-hydrogen-41c.notion.site/Skywork-Open-Reaonser-Series-1d0bc9ae823a80459b46c14...
2025
-
[19]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net, 2021 a . URL h...
2021
-
[20]
Hendrycks, C
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt. Measuring mathematical problem solving with the math dataset. ArXiv preprint, abs/2103.03874, 2021 b . URL https://arxiv.org/abs/2103.03874
2021 arXiv
-
[21]
Hsieh, S
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg. Ruler: What's the real context size of your long-context language models? ArXiv preprint, abs/2404.06654, 2024. URL https://arxiv.org/abs/2404.06654
2024 arXiv
-
[22]
J. Hu, Y. Zhang, Q. Han, D. Jiang, X. Zhang, and H.-Y. Shum. Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model. ArXiv preprint, abs/2503.24290, 2025. URL https://arxiv.org/abs/2503.24290
2025 arXiv
-
[23]
Huang, Y
Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, J. Lei, Y. Fu, M. Sun, and J. He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editor...
2023
-
[24]
International olympiad in informatics, 2024
IOI. International olympiad in informatics, 2024. URL https://ioinformatics.org/
2024
-
[25]
N. Jain, K. Han, A. Gu, W.-D. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. ArXiv preprint, abs/2403.07974, 2024. URL https://arxiv.org/abs/2403.07974
2024 arXiv
-
[26]
Joshi, E
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer. T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension. In R. Barzilay and M.-Y. Kan, editors, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1...
2017 doi
-
[27]
Kwiatkowski, J
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M.-W. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov. Natural questions: A benchmark for question answering resea...
2019 doi
-
[28]
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611--626, 2023
2023
-
[29]
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy. RACE : Large-scale R e A ding comprehension dataset from examinations. In M. Palmer, R. Hwa, and S. Riedel, editors, Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 785--794, Copenhagen...
2017 doi
-
[30]
Leviathan, M
Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii...
2023
-
[31]
H. Li, Y. Zhang, F. Koto, Y. Yang, H. Zhao, Y. Gong, N. Duan, and T. Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese. ArXiv preprint, abs/2306.09212, 2023. URL https://arxiv.org/abs/2306.09212
2023 arXiv
-
[32]
Lightman, V
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe. Let's verify step by step. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 202...
2024
-
[33]
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al. Deepseek-v3 technical report. ArXiv preprint, abs/2412.19437, 2024 a . URL https://arxiv.org/abs/2412.19437
2024 arXiv
-
[34]
J. Liu, C. S. Xia, Y. Wang, and L. Zhang. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Process...
2023
-
[35]
J. Liu, D. Zhu, Z. Bai, Y. He, H. Liao, H. Que, Z. Wang, C. Zhang, G. Zhang, J. Zhang, et al. A comprehensive survey on long context language modeling. ArXiv preprint, abs/2503.17407, 2025. URL https://arxiv.org/abs/2503.17407
2025
-
[36]
X. Liu, X. Lei, S. Wang, Y. Huang, Z. Feng, B. Wen, J. Cheng, P. Ke, Y. Xu, W. L. Tam, X. Zhang, L. Sun, X. Gu, H. Wang, J. Zhang, M. Huang, Y. Dong, and J. Tang. Alignbench: Benchmarking chinese alignment of large language models, 2024 b . URL https://arxiv.org/abs/2311.18743
2024 arXiv
-
[37]
Y. Liu, R. Jin, L. Shi, Z. Yao, and D. Xiong. Finemath: A fine-grained mathematical evaluation benchmark for chinese large language models. ArXiv preprint, abs/2403.07747, 2024 c . URL https://arxiv.org/abs/2403.07747
2024 arXiv
-
[38]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7
2019
-
[39]
American invitational mathematics examination - aime
MAA. American invitational mathematics examination - aime. In American Invitational Mathematics Examination - AIME, 2024. URL https://maa.org/math-competitions/american-invitational-mathematics-examination-aime
2024
-
[40]
American invitational mathematics examination - aime
MAA. American invitational mathematics examination - aime. In American Invitational Mathematics Examination - AIME, 2025. URL https://maa.org/math-competitions/american-invitational-mathematics-examination-aime
2025
-
[41]
Moritz, R
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, et al. Ray: A distributed framework for emerging \ AI \ applications. In 13th USENIX symposium on operating systems design and implementation (OSDI 18), pages 561--577, 2018
2018
-
[42]
Learning to reason with llms, 2024
OpenAI. Learning to reason with llms, 2024. URL https://openai.com/index/learning-to-reason-with-llms/
2024
-
[43]
Paster, M
K. Paster, M. D. Santos, Z. Azerbayev, and J. Ba. Openwebmath: An open dataset of high-quality mathematical web text. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview....
2024
-
[44]
Penedo, Q
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only. ArXiv preprint, abs/2306.01116, 2023. URL https://arxiv....
2023 arXiv
-
[45]
Penedo, H
G. Penedo, H. Kydl \' cek, L. B. Allal, A. Lozhkov, M. Mitchell, C. A. Raffel, L. von Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang, editor...
2024
-
[46]
Radford, K
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. Improving language understanding by generative pre-training. OpenAI, 2018
2018
-
[47]
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First Conference on Language Modeling, 2024
2024
-
[48]
Sakaguchi, R
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi. Winogrande: An adversarial winograd schema challenge at scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IA...
2020
-
[49]
B. Seed, Y. Yuan, Y. Yue, M. Wang, X. Zuo, J. Chen, L. Yan, W. Xu, C. Zhang, X. Liu, et al. Seed-thinking-v1. 5: Advancing superb reasoning models with reinforcement learning. ArXiv preprint, abs/2504.13914, 2025. URL https://arxiv.org/abs/2504.13914
2025
-
[50]
Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. Li, Y. Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. ArXiv preprint, abs/2402.03300, 2024. URL https://arxiv.org/abs/2402.03300
2024 arXiv
-
[51]
Sheng, C
G. Sheng, C. Zhang, Z. Ye, X. Wu, W. Zhang, R. Zhang, Y. Peng, H. Lin, and C. Wu. Hybridflow: A flexible and efficient rlhf framework. ArXiv preprint, abs/2409.19256, 2024. URL https://arxiv.org/abs/2409.19256
2024 arXiv
-
[52]
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568: 0 127063, 2024
2024
-
[53]
Suzgun, N
M. Suzgun, N. Scales, N. Sch \"a rli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. Le, E. Chi, D. Zhou, and J. Wei. Challenging BIG -bench tasks and whether chain-of-thought can solve them. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors, Findings of the Associatio...
2023 doi
-
[54]
C. Team, Z. Yue, Z. Lin, Y. Song, W. Wang, S. Ren, S. Gu, S. Li, P. Li, L. Zhao, L. Li, K. Bao, H. Tian, H. Zhang, G. Wang, D. Zhu, Cici, C. He, B. Ye, B. Shen, Z. Zhang, Z. Jiang, Z. Zheng, Z. Song, Z. Luo, Y. Yu, Y. Wang, Y. Tian, Y. Tu, Y. Yan, Y. Huang, X. Wang, X. Xu, X. ...
2025 arXiv
-
[55]
G. Team. Gemma 2: Improving open language models at a practical size, 2024. URL https://arxiv.org/abs/2408.00118
2024 arXiv
-
[56]
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al. Kimi k1. 5: Scaling reinforcement learning with llms. ArXiv preprint, abs/2501.12599, 2025 b . URL https://arxiv.org/abs/2501.12599
2025 arXiv
-
[57]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, abs/2307.09288, 2023. URL https://arxiv.org/abs/2307.09288
2023 arXiv
-
[58]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Pro...
2017
-
[59]
Y. Wang, X. Ma, G. Zhang, Y. Ni, A. Chandra, S. Guo, W. Ren, A. Arulraj, X. He, Z. Jiang, T. Li, M. Ku, K. Wang, A. Zhuang, R. Fan, X. Yue, and W. Chen. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. In A. Globersons, L. Mackey, D. Belgrav...
2024
-
[60]
H. Xia, T. Ge, P. Wang, S.-Q. Chen, F. Wei, and Z. Sui. Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation. In H. Bouamor, J. Pino, and K. Bali, editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3909--...
2023 doi
-
[61]
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al. Qwen2. 5 technical report. ArXiv preprint, abs/2412.15115, 2024. URL https://arxiv.org/abs/2412.15115
2024 arXiv
-
[62]
Q. Yu, Z. Zhang, R. Zhu, Y. Yuan, X. Zuo, Y. Yue, T. Fan, G. Liu, L. Liu, X. Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. ArXiv preprint, abs/2503.14476, 2025. URL https://arxiv.org/abs/2503.14476
2025 arXiv
-
[63]
Y. Yue, Z. Chen, R. Lu, A. Zhao, Z. Wang, Y. Yue, S. Song, and G. Huang. Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025. URL https://arxiv.org/abs/2504.13837
2025 arXiv
-
[64]
Zellers, A
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. H ella S wag: Can a machine really finish your sentence? In A. Korhonen, D. Traum, and L. M \`a rquez, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4800,...
2019 doi
-
[65]
Zhang and R
B. Zhang and R. Sennrich. Root mean square layer normalization. In H. M. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch \' e - Buc, E. B. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing S...
2019
-
[66]
Zhong, R
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan. AGIE val: A human-centric benchmark for evaluating foundation models. In K. Duh, H. Gomez, and S. Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages ...
2024
-
[67]
Zhong, Z
Y. Zhong, Z. Zhang, B. Wu, S. Liu, Y. Chen, C. Wan, H. Hu, L. Xia, R. Ming, Y. Zhu, et al. Rlhfuse: Efficient rlhf training for large language models with inter-and intra-stage fusion. ArXiv preprint, abs/2409.13221, 2024 b . URL https://arxiv.org/abs/2409.13221
2024 arXiv
-
[68]
F. Zhou, Z. Wang, N. Ranjan, Z. Cheng, L. Tang, G. He, Z. Liu, and E. P. Xing. Megamath: Pushing the limits of open math corpora. ArXiv preprint, abs/2504.02807, 2025. URL https://arxiv.org/abs/2504.02807
2025 arXiv
-
[69]
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou. Instruction-following evaluation for large language models, 2023. URL https://arxiv.org/abs/2311.07911
2023 arXiv
-
[70]
Q. Zhu, D. Guo, Z. Shao, D. Yang, P. Wang, R. Xu, Y. Wu, Y. Li, H. Gao, S. Ma, et al. Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence. ArXiv preprint, abs/2406.11931, 2024. URL https://arxiv.org/abs/2406.11931
2024 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.