REVIEW 5 major objections 5 minor 1 cited by
YuLan-Mini: An Open Data-efficient Language Model
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A 2.42B-parameter model trained on 1.08T tokens matches small-model rivals trained on up to 18T tokens.
desk verdict A transparent, practical training recipe for a 2.4B model, but the headline data-efficiency claim is weakened by benchmark-tuned data decisions and cited baselines. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the three-part training recipe rather than any single technique. (1) The data pipeline combines MinHash de-duplication, heuristic filters, topic classifiers for math/code/reasoning recall, model-based quality scoring, n-gram decontamination, and large-scale synthetic generation of reasoning documents, chain-of-thought solutions, formal Lean proofs, and reflection data; these streams are arranged into 27 curriculum phases of 40B tokens each under a WSD schedule (10B warmup, 990B stable, 80B annealing) with per-phase mixture shifts kept under 3 percent. (2) The stability package scales initialization to $\sigma_{\mathrm{base}} = \sqrt{2/(5d)}$, scales the embedding output by 10, scales residual branches by $1.4\sqrt{n_{\mathrm{layers}}}$, applies µParameterization-style learning-rate scaling to QKV and FFN weights, and adds WeSaR reparameterization $W = \alpha \tilde{W}$ to decouple gradient size from direction, allowing a global learning rate of 0.01 with z-loss and a reduced Adam epsilon. (3) The annealing stage spends 80B tokens on a high-value mix selected by an accelerated gradient-based method (a LESS variant with InsTag), decays the learning rate with a 1-sqrt curve, and raises the RoPE base frequency from 10,000 to 490,000 to extend the context window to 28K tokens while using masked cross-document attention to preserve short-text performance.
What would settle it
Retrain YuLan-Mini from the released checkpoints and per-phase data under a pre-registered protocol — a single fixed chain-of-thought prompt, no curriculum feedback from the test sets, and baselines re-run in the same harness — and check whether the eight-benchmark average still beats Qwen2.5-1.5B and SmolLM2; if it falls behind when prompt choice and data-ratio tuning are taken away, the data-efficiency claim is not robust.
Extended reading notes
Core claim
YuLan-Mini is a 2.42B-parameter model with 56 layers, a 1,920-dimensional hidden width, grouped-query attention, and a 99K-vocabulary tokenizer with digit splitting. Trained on 1.08T tokens, the 28K-context checkpoint scores 37.80 on MATH-500 (4-shot), 64.00 on HumanEval (0-shot), 68.46 on GSM8K, and 49.10 on MMLU (5-shot). The authors argue that the model's aggregate performance, averaged over eight benchmarks, is competitive with small industry models trained on 2–18T tokens, despite using roughly one-half to one-seventeenth of their training budgets. The paper attributes this efficiency to a data pipeline that combines cleaning, classifier-based recall, decontamination, and a 27-phase WSD curriculum; to a stability package built on µParameterization-style initialization plus WeSaR reparameterization; and to an annealing stage that mixes high-value reasoning data with long-context training. It releases the full per-phase token composition to make the whole recipe reproducible.
Load-bearing premise
The entire data-efficiency comparison rests on the evaluation being a fair apples-to-apples contest: baseline scores come from each model's own paper, the better of two chain-of-thought prompts is picked per model, and the training-data mixture is adjusted based on the very benchmarks used in the final ranking, so if those numbers are not directly comparable the claimed token savings could shrink.
Editorial extensions
If this is right
- A 2.42B base model can reach the top of its size class on math and code benchmarks after 1.08T tokens, with the 28K checkpoint scoring 37.80 on MATH-500, 68.46 on GSM8K, and 64.00 on HumanEval.
- The 27-phase WSD curriculum, with per-phase mixture shifts capped at 3%, keeps training stable across 1.08T tokens at a global learning rate of 0.01.
- The annealing stage extends the context window from 4K to 28K tokens by raising the RoPE base frequency to 490,000 while using long-context data with masked cross-document attention to preserve short-text skills.
- Because the full per-phase data composition is released, the recipe can be reproduced and its components (data selection, annealing mix, stability package) can be ablated by other groups.
Reading between the lines
- Whether the same 1.08T-token budget suffices at 7B scale is an open question; the stability package is designed to transfer via µParameterization, but the data curriculum and annealing are tuned for 2.4B.
- The evaluation protocol's prompt selection (better of two per model) could inflate YuLan-Mini's margin; a single pre-registered prompt would make the data-efficiency claim sharper.
- The annealing mix blends long-thought, formal-math, and gradient-selected data; isolating these components would show which one drives the MATH-500 and HumanEval gains.
- The 4K and 28K checkpoints trade off general knowledge (MMLU 51.79 vs 49.10) for reasoning gains (MATH-500 32.60 vs 37.80); deployment choices may favor different checkpoints.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports the pre-training of YuLan-Mini, a 2.42B-parameter decoder-only base model trained on about 1.08T tokens, and claims that it reaches performance comparable to industry models trained on substantially more data. The technical recipe has three pillars: a data pipeline with cleaning, mixing, and curriculum scheduling; a stability-focused optimization setup combining scaled initialization, µP-like rules, and WeSaR re-parameterization; and an annealing stage with targeted data selection, learning-rate annealing, and context extension to 28K. The authors evaluate the model on a suite of math, code, commonsense, and Chinese benchmarks, and release detailed per-phase data compositions and checkpoints. The central claim is that the reported benchmark results demonstrate a data-efficient pre-training recipe reproducible in a university setting.
Significance. If the data-efficiency claim holds, the paper is a valuable resource for the community: it provides unusually detailed disclosure of data composition at each curriculum phase, a systematic discussion of training-stability diagnostics and mitigations, and an open release that lowers the barrier for reproducing competitive small models. The per-phase data tables and the stability analysis are concrete contributions that do not depend on the contested comparison. However, the headline claim depends on the evaluation protocol being a fair and controlled comparison, and that premise is weakened by the paper's own description of how the data mixture and annealing data were chosen.
major comments (5)
- [Section 4.5, Section 5.2, Tables 6-7] The central data-efficiency claim is not supported by a controlled evaluation protocol. Section 4.5 states that at each 40B-token curriculum boundary the data ratios are reassessed and adjusted based on the model's overall performance, with HumanEval given as an explicit example, and Section 5.2 states that formal math and o1-like reasoning data were incorporated to improve performance on challenging math benchmarks such as MATH-500. These are the same benchmarks used in Tables 6-7 and Figure 1 to argue for superior data efficiency. The reported math/code advantage may therefore reflect iterative tuning of the data mixture against the evaluation set rather than a generally more data-efficient pre-training recipe. I ask the authors to either hold out a set of benchmarks that were never used for data-mix decisions, report the trajectory of decisions together with the resulting scores, or evaluate frozen checkpoints from the originally scheduled curriculum.
- [Section 6.1.3, Tables 6-7] The comparison against baselines is not made under identical conditions. The text states that for CoT benchmarks the authors evaluate each model with both a short and a long prompt and select the higher score, while baseline numbers are mostly cited from official reports rather than re-run in the same harness (the table marks several values with an asterisk). This per-model prompt selection, combined with heterogeneous evaluation sources, can systematically inflate the relative standing of YuLan-Mini. The authors should re-evaluate all baselines with the identical evaluation code, generation limits, and a single fixed prompt per task, and report both prompt variants or justify why prompt selection cannot favor their model.
- [Section 5.1, Section 2.4] The annealing ratio and annealing function are fitted choices whose selection is part of the final model. The paper estimates the 8% annealing ratio from a scaling law and states that 1-sqrt annealing was chosen because it performed best empirically, and the final model is produced with that exact configuration. Since the headline comparison is made with the final model only, the reader cannot separate the effect of the annealing strategy from the effect of having selected the best-performing configuration on the evaluation benchmarks. At minimum, the paper should report the performance of checkpoints before annealing and, if available, results from alternative annealing ratios or functions.
- [Section 2.5 and Table 1] The paper's stability story relies partly on proxy-model experiments, but the transfer of the stability conclusions from the 0.05B/0.2B proxy models to the 2.42B model is asserted rather than demonstrated. Section 3.2.2 mentions that instability still appeared when migrating to the target size, and the mitigation is then validated mainly on small proxies and on the final successful run. Given that training-stability claims are one of the three advertised contributions, the authors should provide at least a controlled comparison showing that the chosen initialization/re-parameterization combination prevents divergence on the target-scale model under conditions where the baseline diverges.
- [Table 6 and Section 2.3] There is a discrepancy in the token count used for the main claim. The abstract and Section 2.3 state 1.08T tokens, but Table 6 reports the 4K checkpoint as trained on 1.04T tokens and the 28K checkpoint as trained on 1.08T tokens. Since the data-efficiency comparison in Figure 1 uses the average scores of the final model, the exact token count attributed to the evaluated checkpoint must be clarified. If the 4K checkpoint is the one used for parts of the comparison, the paper should state whether the comparison uses the 1.04T or 1.08T checkpoint.
minor comments (5)
- [Throughout] There are several typographical and grammatical errors: 'diffrent' in Table 1, 'intergration' in Section 2.5, 'have have' in Section 3.3.2, and 'hightlited' in Section 4.2. These should be corrected.
- [Section 2.4 and Table 3] The residual connection scaling factor is listed as 1.4√n_layers in Table 3 for YuLan-Mini, but the text in Section 3.2 does not derive or motivate this specific value; please add a short explanation or reference.
- [Figure 1] The caption says that models larger than 3B are plotted in gray, but the legend and axis labels do not make it easy to identify which points correspond to which models; please add labels to the points or a clearer legend.
- [Section 6.1.3] The evaluation section states that gpt-4o-mini is used to verify MATH-500 outputs and manual checks were conducted, but the number of samples checked manually is not given; please specify the verification procedure and sample size.
- [Appendix E] The detailed phase tables are useful, but the row for Phase 1 says the first 10B tokens are warmup and the next 30B are stable training; this should be stated directly above the table as well as in the main text for readability.
Circularity Check
Data-efficiency claim is partially circular: the data curriculum and annealing mix were tuned against the same benchmarks (HumanEval, MATH-500, GSM8K) that are then reported as evidence of superior data efficiency.
-
fitted input called prediction
[Section 4.5 (Data Curriculum), p.18; evidence used in Section 6.2 and Tables 6-7.]
"For each 40B tokens, we reassess and adjust the data ratio when transitioning between training phases based on the model's overall performance in that phase. For example, if the model's performance on the HumanEval benchmark does not improve or declines after a stage, we may consider slightly increasing the amount of code data in the subsequent stage."
The paper's headline claim of 'superior data efficiency' is supported by average scores on GSM8K, MATH-500, HumanEval, MBPP, MMLU, ARC-Challenge, HellaSwag, and CEval (Figure 1, Tables 6-7). Section 4.5 states that the data mixture itself was reassessed and adjusted every 40B tokens using the model's performance on benchmarks such as HumanEval. The final scores on those benchmarks are therefore partly optimized targets of the curriculum, not independent measurements of a fixed pretraining recipe. Reporting them as evidence that 1.08T tokens suffice is a fitted-input-called-prediction loop: the recipe was fit to the evaluation set and then evaluated on the same set.
-
fitted input called prediction
[Section 5.2 (Data Selection for Annealing Stage), p.19; compared with Table 6 MATH-500 row.]
"In particular, we incorporate formal mathematical reasoning (theorem proving in Lean) and advanced reasoning data (o1-like thought data) to improve the model's performance on challenging math benchmarks, e.g., MATH-500, which have been shown in Table 6."
This sentence explicitly states that annealing-stage data were selected with the goal of improving MATH-500, and then points to Table 6, where YuLan-Mini's MATH-500 score (37.80) is reported as evidence of mathematical capability and of training efficacy. The MATH-500 gain is the objective of the data-selection procedure, so citing it as an independent confirmation of the annealing approach is circular: the benchmark improvement is what the selection was engineered to produce, not a predicted consequence of a general recipe.
full rationale
The paper is transparent and mostly self-contained about architecture, data, and optimization choices; the training-stability analysis in Section 3 is an independent empirical study with proxy models, and the annealing-ratio estimate (8%) is taken from an external scaling law (Tissue et al., 2024), which is not circular. However, the central data-efficiency claim is weakened by a genuine feedback loop that the paper itself documents. Section 4.5 says data ratios were reassessed and adjusted at each 40B boundary based on model performance on benchmarks, with HumanEval as the example; Section 5.2 says annealing data were chosen to improve MATH-500. The same benchmarks are then used in Tables 6-7 and Figure 1 to argue for 'superior training efficacy' from only 1.08T tokens. Consequently, the reported math/code advantages are at least partially optimized targets rather than independent evidence of a generally data-efficient recipe. The evaluation section also acknowledges a second limitation: baseline scores are cited from official reports rather than re-run under the identical protocol ('fully reproducing the results of these baseline models as originally reported remains challenging'), so the comparison does not control for this tuning loop. This is an experimental-design confound, not an allegation of misconduct, and it does not affect the paper's other contributions such as the stability methods, data release, and curriculum transparency. Because two of the headline 'predictions' (data efficiency on math/code benchmarks) reduce to benchmark-driven fitting, a score of 6 is appropriate.
Assumptions & free parameters
free parameters (8)
- Learning rate =
0.01
- Warmup tokens =
10B
- Annealing ratio =
8% (80B tokens)
- RoPE base frequency (annealing) =
490,000
- Embedding scaling factor =
10
- Data mixture proportions =
60% English, 20% code, 10% math, 10% Chinese, with per-phase adjustments up to 3%
- BPE dropout rate =
0.2
- Z-loss coefficient =
1e-4
assumptions (6)
- standard math Transformer architecture and next-token prediction objective are effective for language modeling.
- domain assumption Scaling laws (Kaplan et al.) accurately estimate FLOPs and annealing behavior.
- domain assumption Benchmark scores from different papers are comparable when evaluated under similar settings.
- domain assumption Synthetic data from Qwen2.5-Math and QwQ-32B-Preview does not introduce benchmark contamination beyond the n-gram decontamination.
- ad hoc to paper Hidden states variance and gradient norm are reliable indicators of training instability that transfer from 0.2B proxy models to 2.42B model.
- ad hoc to paper The 1-sqrt annealing function outperforms alternatives.
Cite this review
Pith. "Pith review of YuLan-Mini: An Open Data-efficient Language Model." pith.science (2026). https://pith.science/paper/H2ZKK26I
@misc{pith2026241217743,
author = {Pith},
title = {Pith review of: YuLan-Mini: An Open Data-efficient Language Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/H2ZKK26I}},
note = {Machine review of arXiv:2412.17743}
}
read the original abstract
Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the complexity of the technical processes involved. This paper presents a detailed technical report on YuLan-Mini, a highly capable base model with 2.42B parameters that achieves top-tier performance among models of similar parameter scale. Our pre-training approach focuses on enhancing training efficacy through three key technical contributions: an elaborate data pipeline combines data cleaning with data schedule strategies, a robust optimization method to mitigate training instability, and an effective annealing approach that incorporates targeted data selection and long context training. Remarkably, YuLan-Mini, trained on 1.08T tokens, achieves performance comparable to industry-leading models that require significantly more data. To facilitate reproduction, we release the full details of the data composition for each training phase. Project details can be accessed at the following link: https://github.com/RUC-GSAI/YuLan-Mini.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance
Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.
Reference graph
Works this paper leans on
-
[1]
Synthetically generated reasoning dataset (gsm8k-inspired) with enhanced diversity using gretel navigator and meta-llama/meta-llama-3.1-405b
Gretel AI. Synthetically generated reasoning dataset (gsm8k-inspired) with enhanced diversity using gretel navigator and meta-llama/meta-llama-3.1-405b. https://huggingface.co/gretelai/synthetic-gsm8k-reflection-405b, 9 2024
2024
-
[2]
GQA: training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee - Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebr \' o n, and Sumit Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singa...
-
[3]
Smollm2 - with great data, comes great performance, 2024
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Lewis Tunstall, Agustín Piqueres, Andres Marafioti, Cyril Zakka, Leandro von Werra, and Thomas Wolf. Smollm2 - with great data, comes great performance, 2024
2024
-
[4]
Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021. URL https://arxiv.org/abs/2108.07732
arXiv 2021
-
[5]
Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://open...
2024
-
[6]
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. CoRR, abs/1607.06450, 2016. URL http://arxiv.org/abs/1607.06450
arXiv 2016
-
[7]
Numinamath 7b cot, 2024
Edward Beeching, Shengyi Costa Huang, Albert Jiang, Jia Li, Benjamin Lipkin, Zihan Qina, Kashif Rasul, Ziju Shen, Roman Soletskyi, and Lewis Tunstall. Numinamath 7b cot, 2024. URL http://faculty.bicmr.pku.edu.cn/ dongbin/Publications/numina_dataset.pdf
2024
-
[8]
Stable LM 2 1.6b technical report
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparaju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccol \' o Zanichelli, and Carlos Riquelme. Stable LM 2 1.6b technical report. CoRR, abs/2402.17834...
Show all 135 references
- [9]
-
[10]
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5: 0 135--146, 2017. ISSN 2307-387X
2017
-
[11]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020 arXiv
-
[12]
Towards effective and efficient continual pre-training of large language models
Jie Chen, Zhipeng Chen, Jiapeng Wang, Kun Zhou, Yutao Zhu, Jinhao Jiang, Yingqian Min, Wayne Xin Zhao, Zhicheng Dou, Jiaxin Mao, Yankai Lin, Ruihua Song, Jun Xu, Xu Chen, Rui Yan, Zhewei Wei, Di Hu, Wenbing Huang, and Ji - Rong Wen. Towards effective and efficient continual pr...
-
[13]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Sco...
2021 arXiv
-
[14]
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending context window of large language models via positional interpolation. CoRR, abs/2306.15595, 2023. doi:10.48550/ARXIV.2306.15595. URL https://doi.org/10.48550/arXiv.2306.15595
-
[15]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...
2023
-
[16]
Stable language model pre-training by reducing embedding variability
Woojin Chung, Jiwoo Hong, Na Min An, James Thorne, and Se - Young Yun. Stable language model pre-training by reducing embedding variability. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural La...
2024
-
[17]
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://...
-
[18]
Getting the most out of your tokenizer for pre-training and domain adaptation
Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozi \` e re. Getting the most out of your tokenizer for pre-training and domain adaptation. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL http...
2024
-
[19]
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=mZn2Xyh9Ec
2024
-
[20]
Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e . Flashattention: Fast and memory-efficient exact attention with io-awareness. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Proces...
2022
-
[21]
The z-loss: a shift and scale invariant classification loss belonging to the spherical family
Alexandre de Br \' e bisson and Pascal Vincent. The z-loss: a shift and scale invariant classification loss belonging to the spherical family. CoRR, abs/1604.08859, 2016. URL http://arxiv.org/abs/1604.08859
2016 arXiv
- [22]
-
[23]
Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster
Nolan Dey, Gurpreet Gosal, Zhiming Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster. CoRR, abs/2304.03208, 2023 a . doi:10.48550/ARXIV.2304.0320...
-
[24]
Cerebras- GPT : Open Compute - Optimal Language Models Trained on the Cerebras Wafer - Scale Cluster , April 2023 b
Nolan Dey, Gurpreet Gosal, Zhiming, Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness. Cerebras- GPT : Open Compute - Optimal Language Models Trained on the Cerebras Wafer - Scale Cluster , April 2023 b . URL http://arxiv.org/abs/2304.03208....
2023 arXiv
-
[25]
Fewer truncations improve language modeling
Hantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar, Anoop Deoras, Dan Roth, and Stefano Soatto. Fewer truncations improve language modeling. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024 a...
2024
-
[26]
Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , October 2024 b
Yuyang Ding, Xinyu Shi, Xiaobo Liang, Juntao Li, Qiaoming Zhu, and Min Zhang. Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , October 2024 b . URL http://arxiv.org/abs/2410.18693. arXiv:2410.18693 [cs]
2024 arXiv
- [27]
-
[28]
How to train long-context language models (effectively)
Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen. How to train long-context language models (effectively). CoRR, abs/2410.02660, 2024. doi:10.48550/ARXIV.2410.02660. URL https://doi.org/10.48550/arXiv.2410.02660
2024 doi
-
[29]
Dirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu,...
2024
- [30]
-
[31]
Scaling laws and compute-optimal training beyond fixed training durations
Alexander H \" a gele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro von Werra, and Martin Jaggi. Scaling laws and compute-optimal training beyond fixed training durations. CoRR, abs/2405.18392, 2024. doi:10.48550/ARXIV.2405.18392. URL https://doi.org/10.48550/arXiv.2405.18392
-
[32]
Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning
Xiaotian Han, Yiren Jian, Xuefeng Hu, Haogeng Liu, Yiqi Wang, Qihang Fan, Yuang Ai, Huaibo Huang, Ran He, Zhenheng Yang, and Quanzeng You. Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning. CoRR, abs/2409.12568, 2024. doi:10.48550/ARXIV....
-
[33]
Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models
Conghui He, Zhenjiang Jin, Chao Xu, Jiantao Qiu, Bin Wang, Wei Li, Hang Yan, Jiaqi Wang, and Dahua Lin. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models. CoRR, abs/2308.10755, 2023. doi:10.48550/ARXIV.2308.10755. URL https://doi.org/10...
-
[34]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview...
2021
-
[35]
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Joaquin Vanschoren and Sai - Kit Yeung, editors, Proceedings of the Neural Information Processi...
2021
-
[36]
RULER: what's the real context size of your long-context language models? CoRR, abs/2404.06654, 2024
Cheng - Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. RULER: what's the real context size of your long-context language models? CoRR, abs/2404.06654, 2024. doi:10.48550/ARXIV.2404.06654. URL https://doi.org/10.48...
-
[37]
Liger kernel: Efficient triton kernels for LLM training
Pin - Lun Hsu, Yun Dai, Vignesh Kothapalli, Qingquan Song, Shao Tang, Siyu Zhu, Steven Shimizu, Shivam Sahni, Haowen Ning, and Yanning Chen. Liger kernel: Efficient triton kernels for LLM training. CoRR, abs/2410.10989, 2024. doi:10.48550/ARXIV.2410.10989. URL https://doi.org/...
-
[38]
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zhen Leng Thai, Kai Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai Li, Z...
-
[39]
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen - Chuan Chang. Towards reasoning in large language models: A survey. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 104...
2023 doi
- [40]
-
[41]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural ...
2023
-
[42]
Technical report: Enhancing llm reasoning with reward-guided tree search
Jinhao Jiang, Zhipeng Chen, Yingqian Min, Jie Chen, Xiaoxue Cheng, Jiapeng Wang, Yiru Tang, Haoxiang Sun, Jia Deng, Wayne Xin Zhao, Zheng Liu, Dong Yan, Jian Xie, Zhongyuan Wang, and Ji-Rong Wen. Technical report: Enhancing llm reasoning with reward-guided tree search. CoRR, a...
-
[43]
TinyBERT : Distilling BERT for Natural Language Understanding , October 2020
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. TinyBERT : Distilling BERT for Natural Language Understanding , October 2020. URL http://arxiv.org/abs/1909.10351. Issue: arXiv:1909.10351 1097 citations (Semantic Scholar/arXiv) [2...
2020 arXiv
-
[44]
Calc-x and calcformers: Empowering arithmetical chain-of-thought through interaction with symbolic systems
Marek Kadlc \' k, Michal Stef \' a nik, Ondrej Sotol \' a r, and Vlastimil Martinek. Calc-x and calcformers: Empowering arithmetical chain-of-thought through interaction with symbolic systems. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Confe...
2023 doi
-
[45]
Kaplan, Sam McCandlish, T
J. Kaplan, Sam McCandlish, T. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. Scaling Laws for Neural Language Models . ArXiv, January 2020. URL https://www.semanticscholar.org/paper/Scaling-Laws-for-Neural-Language-Mod...
2020
-
[46]
LAMBADA: backward chaining for automated reasoning in natural language
Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, and Deepak Ramachandran. LAMBADA: backward chaining for automated reasoning in natural language. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association f...
2023 doi
-
[47]
Strategic data ordering: Enhancing large language model performance through curriculum learning
Jisu Kim and Juhwan Lee. Strategic data ordering: Enhancing large language model performance through curriculum learning. CoRR, abs/2405.07490, 2024. doi:10.48550/ARXIV.2405.07490. URL https://doi.org/10.48550/arXiv.2405.07490
-
[48]
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Jason Flinn, Margo I. Seltzer, Peter Druschel, Antoine Kaufmann, and...
2023
-
[49]
Race: Large-scale reading comprehension dataset from examinations, 2017
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations, 2017. URL https://arxiv.org/abs/1704.04683
2017 arXiv
-
[50]
Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...
-
[51]
To FP8 and back again: Quantifying the effects of reducing precision on LLM training stability
Joonhyung Lee, Jeongin Bae, Byeongwook Kim, Se Jung Kwon, and Dongsoo Lee. To FP8 and back again: Quantifying the effects of reducing precision on LLM training stability. CoRR, abs/2405.18710, 2024. doi:10.48550/ARXIV.2405.18710. URL https://doi.org/10.48550/arXiv.2405.18710
-
[52]
The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models
Conglong Li, Minjia Zhang, and Yuxiong He. The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems ...
2022
-
[53]
CMMLU: measuring massive multitask language understanding in chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. CMMLU: measuring massive multitask language understanding in chinese. In Lun - Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computationa...
2024 doi
-
[54]
Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Raghavi Chandu, Thao Nguyen, Igor Vasiljevic, Sham M
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Alb...
-
[55]
Numinamath
Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu. Numinamath. [https://github.com/project-numina/aimo-...
2024
-
[56]
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy - Poirier, Jo \ a o Mont...
2023
-
[57]
Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023
Wing Lian, Guan Wang, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023. URL https://https://huggingface.co/Open-Orca/SlimOrca
2023
-
[58]
Universal checkpointing: Efficient and flexible checkpointing for large scale distributed training
Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase, and Minjia Zhang. Universal checkpointing: Efficient and flexible checkpointing for large scale distributed training. CoRR, abs/2406.18820, 2024. doi:10.48550/ARXIV.2406.18820. URL https:...
-
[59]
Let's verify step by step
Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let's verify step by step. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-1...
2024
- [60]
-
[61]
Evaluating language models for efficient code generation
Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. Evaluating language models for efficient code generation. In First Conference on Language Modeling, 2024 a . URL https://openreview.net/forum?id=IBCBMeAhmC
2024
-
[62]
Longwanjuan: Towards systematic measurement for long text quality
Xiaoran Liu, Kai Lv, Qipeng Guo, Hang Yan, Conghui He, Xipeng Qiu, and Dahua Lin. Longwanjuan: Towards systematic measurement for long text quality. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Findings of the Association for Computational Linguistics: EM...
2024
-
[63]
Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra
Zechun Liu, Changsheng Zhao, Forrest N. Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra. Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. I...
2024
-
[64]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7
2019
-
[65]
Fineweb-edu, May 2024 a
Anton Lozhkov, Loubna Ben Allal, Leandro von Werra, and Thomas Wolf. Fineweb-edu, May 2024 a . URL https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu
2024
-
[66]
McAuley, Han Hu, Torsten Scholak, S \' e bastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, and et al
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy - Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul,...
-
[67]
\# instag: Instruction tagging for analyzing supervised fine-tuning of large language models
Keming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin, Junyang Lin, Chuanqi Tan, Chang Zhou, and Jingren Zhou. \# instag: Instruction tagging for analyzing supervised fine-tuning of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, ...
2024
- [68]
-
[69]
Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems
Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, and Ji-Rong Wen. Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning system...
-
[70]
Agentinstruct: Toward generative teaching with agentic flows
Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andr \' e s Codas, Yadong Lu, Weige Chen, Olga Vrousgos, Corby Rosset, Fillipe Silva, Hamed Khanpour, Yash Lara, and Ahmed Awadallah. Agentinstruct: Toward generative teaching with agentic flows. Co...
-
[71]
Orca-math: Unlocking the potential of slms in grade school math
Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. Orca-math: Unlocking the potential of slms in grade school math. CoRR, abs/2402.14830, 2024 b . doi:10.48550/ARXIV.2402.14830. URL https://doi.org/10.48550/arXiv.2402.14830
-
[72]
A theory on adam instability in large-scale machine learning
Igor Molybog, Peter Albert, Moya Chen, Zachary DeVito, David Esiobu, Naman Goyal, Punit Singh Koura, Sharan Narang, Andrew Poulton, Ruan Silva, Binh Tang, Diana Liskovich, Puxin Xu, Yuchen Zhang, Melanie Kambadur, Stephen Roller, and Susan Zhang. A theory on adam instability i...
-
[73]
A corpus and evaluation framework for deeper understanding of commonsense stories, 2016
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and evaluation framework for deeper understanding of commonsense stories, 2016. URL https://arxiv.org/abs/1604.01696
2016 arXiv
-
[74]
Initialization of large language models via reparameterization to mitigate loss spikes
Kosuke Nishida, Kyosuke Nishida, and Kuniko Saito. Initialization of large language models via reparameterization to mitigate loss spikes. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Lang...
2024
- [75]
-
[76]
Opencsg/chinese-fineweb-edu Datasets at Hugging Face
Opencsg. Opencsg/chinese-fineweb-edu Datasets at Hugging Face . https://huggingface.co/datasets/opencsg/chinese-fineweb-edu
-
[77]
Openwebmath: An open dataset of high-quality mathematical web text
Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. Openwebmath: An open dataset of high-quality mathematical web text. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL htt...
2024
-
[78]
The FineWeb Datasets : Decanting the Web for the Finest Text Data at Scale , October 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The FineWeb Datasets : Decanting the Web for the Finest Text Data at Scale , October 2024. URL http://arxiv.org/abs/2406.17557. arXiv:2406.17557
2024 arXiv
-
[79]
Using the output embedding to improve language models
Ofir Press and Lior Wolf. Using the output embedding to improve language models. In Mirella Lapata, Phil Blunsom, and Alexander Koller, editors, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, Valencia, Sp...
2017 doi
-
[80]
Bpe-dropout: Simple and effective subword regularization
Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. Bpe-dropout: Simple and effective subword regularization. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistic...
2020 doi
-
[81]
Qwen2.5: A party of foundation models, September 2024
Qwen-Team . Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/blog/qwen2.5/
2024
-
[82]
Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Qwen-Team. Qwq: Reflect deeply on the boundaries of the unknown, November 2024. URL https://qwenlm.github.io/blog/qwq-32b-preview/
2024
-
[83]
Zero: memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: memory optimizations toward training trillion parameter models. In Christine Cuicchi, Irene Qualters, and William T. Kramer, editors, Proceedings of the International Conference for High Performance Comput...
2020 arXiv
- [84]
-
[85]
Winogrande: an adversarial winograd schema challenge at scale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: an adversarial winograd schema challenge at scale. Commun. ACM , 64 0 (9): 0 99--106, 2021. doi:10.1145/3474381. URL https://doi.org/10.1145/3474381
2021 doi
-
[86]
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. Analysing mathematical reasoning abilities of neural models. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openr...
2019
-
[87]
Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M. Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Bel...
2022 arXiv
-
[88]
GLU variants improve transformer
Noam Shazeer. GLU variants improve transformer. CoRR, abs/2002.05202, 2020. URL https://arxiv.org/abs/2002.05202
2002 arXiv
-
[89]
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: language agents with verbal reinforcement learning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information...
2023
-
[90]
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020. URL https://arxiv.org/abs/1909.08053
2020 arXiv
-
[91]
Scaling synthetic logical reasoning datasets with context-sensitive declarative grammars
Damien Sileo. Scaling synthetic logical reasoning datasets with context-sensitive declarative grammars. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami,...
2024
-
[92]
Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, N...
2024
-
[93]
An integrated data processing framework for pretraining foundation models
Yiding Sun, Feng Wang, Yutao Zhu, Wayne Xin Zhao, and Jiaxin Mao. An integrated data processing framework for pretraining foundation models. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang, editors, Proceedings of the 47th International ACM...
2024
-
[94]
Spike no more: Stabilizing the pre-training of large language models
Sho Takase, Shun Kiyono, Sosuke Kobayashi, and Jun Suzuki. Spike no more: Stabilizing the pre-training of large language models. CoRR, abs/2312.16903, 2023. doi:10.48550/ARXIV.2312.16903. URL https://doi.org/10.48550/arXiv.2312.16903
-
[95]
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...
2024
-
[96]
LLMBox : A Comprehensive Library for Large Language Models
Tianyi Tang, Hu Yiwen, Bingqian Li, Wenyang Luo, ZiJing Qin, Haoxiang Sun, Jiapeng Wang, Shiyi Xu, Xiaoxue Cheng, Geyang Guo, Han Peng, Bowen Zheng, Yiru Tang, Yingqian Min, Yushuo Chen, Jie Chen, Ranchi Zhao, Luran Ding, Yuhao Wang, Zican Dong, Xia Chunxuan, Junyi Li, Kun Zho...
2024
-
[97]
Mathscale: Scaling instruction tuning for mathematical reasoning
Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. Mathscale: Scaling instruction tuning for mathematical reasoning. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024 c . URL https://openrev...
2024
-
[98]
Gemma Team. Gemma. 2024. doi:10.34740/KAGGLE/M/3301. URL https://www.kaggle.com/m/3301
2024 doi
-
[99]
D4: improving LLM pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos. D4: improving LLM pretraining via document de-duplication and diversification. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information P...
2023
- [100]
-
[101]
Openmathinstruct-1: A 1.8 million math instruction tuning dataset
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. Openmathinstruct-1: A 1.8 million math instruction tuning dataset. CoRR, abs/2402.10176, 2024. doi:10.48550/ARXIV.2402.10176. URL https://doi.org/10.48550/arXiv.2402.10176
-
[102]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie - Anne Lachaux, Timoth \' e e Lacroix, Baptiste Rozi \` e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur \' e lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient fo...
-
[103]
Tokenization matters! degrading large language models through challenging their tokenization
Dixuan Wang, Yanda Li, Junyuan Jiang, Zepeng Ding, Guochao Jiang, Jiaqing Liang, and Deqing Yang. Tokenization matters! degrading large language models through challenging their tokenization. CoRR, abs/2405.17067, 2024 a . doi:10.48550/ARXIV.2405.17067. URL https://doi.org/10....
-
[104]
Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning
Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li. Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning. In The Twelfth International Conference on Learning Representations, ...
2024
-
[105]
How do your code llms perform? empowering code instruction tuning with really good data
Yejie Wang, Keqing He, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, and Weiran Xu. How do your code llms perform? empowering code instruction tuning with really good data. In Yaser A...
2024
-
[106]
Chi, Quoc V
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, ...
2022
-
[107]
Magicoder: Empowering code generation with oss-instruct
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. Magicoder: Empowering code generation with oss-instruct. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL https://openreview...
2024
-
[108]
Liu, Lechao Xiao, Katie E
Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett, Alexander A. Alemi, Ben Adlam, John D. Co - Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl - Dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. Small-scale pr...
2024
-
[109]
Rabe, Wenda Li, Jimmy Ba, Roger B
Yuhuai Wu, Markus N. Rabe, Wenda Li, Jimmy Ba, Roger B. Grosse, and Christian Szegedy. LIME: learning inductive bias for primitives of mathematical reasoning. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 20...
2021
-
[110]
Lean-github: Compiling github LEAN repositories for a versatile LEAN prover
Zijian Wu, Jiayu Wang, Dahua Lin, and Kai Chen. Lean-github: Compiling github LEAN repositories for a versatile LEAN prover. CoRR, abs/2407.17227, 2024. doi:10.48550/ARXIV.2407.17227. URL https://doi.org/10.48550/arXiv.2407.17227
-
[111]
LESS: selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. LESS: selecting influential data for targeted instruction tuning. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. ...
2024
-
[112]
Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data
Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, and Xiaodan Liang. Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data. CoRR, abs/2405.14333, 2024. doi:10.48550/ARXIV.2405.14333. URL https://doi.org/1...
-
[113]
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie - Yan Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 ...
2020
-
[114]
Effective long-context scaling of foundation models
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, S...
2024
-
[115]
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing. CoRR, abs/2406.08464, 2024. doi:10.48550/ARXIV.2406.08464. URL https://doi.org/10.485...
-
[116]
Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering
Vikas Yadav, Steven Bethard, and Mihai Surdeanu. Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Met...
2019
-
[117]
Baichuan 2: Open large-scale language models
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, Juntao Dai, Kun Fa...
-
[118]
Qwen2 technical report
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...
-
[119]
Qwen2.5-math technical report: Toward mathematical expert model via self-improvement
An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. Qwen2.5-math technical report: Toward mathematical expert model via se...
-
[120]
Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs V: tuning large neural networks via zero-shot hyperparameter transfer. CoRR, abs/2203.03466, 2022. doi:10.48550/ARXIV.2...
-
[121]
Tensor programs VI: feature learning in infinite depth neural networks
Greg Yang, Dingli Yu, Chen Zhu, and Soufiane Hayou. Tensor programs VI: feature learning in infinite depth neural networks. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024 c . URL https://op...
2024
-
[122]
Lean workbook: A large-scale lean problem set formalized from natural language math problems
Huaiyuan Ying, Zijian Wu, Yihan Geng, Jiayu Wang, Dahua Lin, and Kai Chen. Lean workbook: A large-scale lean problem set formalized from natural language math problems. CoRR, abs/2406.03847, 2024. doi:10.48550/ARXIV.2406.03847. URL https://doi.org/10.48550/arXiv.2406.03847
-
[123]
Yoo, Morris A
Andy B. Yoo, Morris A. Jette, and Mark Grondona. SLURM: simple linux utility for resource management. In Dror G. Feitelson, Larry Rudolph, and Uwe Schwiegelshohn, editors, Job Scheduling Strategies for Parallel Processing, 9th International Workshop, JSSPP 2003, Seattle, WA, U...
2003 doi
-
[124]
Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representation...
2024
-
[125]
Mammoth: Building math generalist models through hybrid instruction tuning
Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mammoth: Building math generalist models through hybrid instruction tuning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 ....
2024
-
[126]
Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Llu \' s M \` a rquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguisti...
2019 doi
-
[127]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch \' e - Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural ...
2019
- [128]
-
[129]
Automathtext: Autonomous data selection with language models for mathematical texts
Yifan Zhang, Yifan Luo, Yang Yuan, and Andrew Chi - Chih Yao. Automathtext: Autonomous data selection with language models for mathematical texts. CoRR, abs/2402.07625, 2024 b . doi:10.48550/ARXIV.2402.07625. URL https://doi.org/10.48550/arXiv.2402.07625
-
[130]
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian - Yun Nie, and Ji - Ro...
-
[131]
Opencodeinterpreter: Integrating code generation with execution and refinement
Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. Opencodeinterpreter: Integrating code generation with execution and refinement. In Lun - Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for C...
2024 doi
-
[132]
Programming every example: Lifting pre-training data quality like experts at scale
Fan Zhou, Zengzhi Wang, Qian Liu, Junlong Li, and Pengfei Liu. Programming every example: Lifting pre-training data quality like experts at scale. CoRR, abs/2409.17115, 2024 a . doi:10.48550/ARXIV.2409.17115. URL https://doi.org/10.48550/arXiv.2409.17115
-
[133]
Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models
Kun Zhou, Beichen Zhang, Jiapeng Wang, Zhipeng Chen, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang, and Ji - Rong Wen. Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models. CoRR, abs/2405.14365, 2024 b . doi:10.48550/ARXIV.24...
-
[134]
Yulan: An open-source large language model
Yutao Zhu, Kun Zhou, Kelong Mao, Wentong Chen, Yiding Sun, Zhipeng Chen, Qian Cao, Yihan Wu, Yushuo Chen, Feng Wang, Lei Zhang, Junyi Li, Xiaolei Wang, Lei Wang, Beichen Zhang, Zican Dong, Xiaoxue Cheng, Yuhan Chen, Xinyu Tang, Yupeng Hou, Qiangqiang Ren, Xincheng Pang, Shufan...
-
[135]
Designing effective sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Designing effective sparse expert models. CoRR, abs/2202.08906, 2022. URL https://arxiv.org/abs/2202.08906
2022 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.