Pith. sign in

REVIEW 5 major objections 5 minor 1 cited by

YuLan-Mini: An Open Data-efficient Language Model

T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A 2.42B-parameter model trained on 1.08T tokens matches small-model rivals trained on up to 18T tokens.

desk verdict A transparent, practical training recipe for a 2.4B model, but the headline data-efficiency claim is weakened by benchmark-tuned data decisions and cited baselines. read the letter →

arxiv 2412.17743 v2 pith:H2ZKK26I submitted 2024-12-23 cs.CL

classification cs.CL
keywords data-efficientpre-trainingsmalllanguagemodelstrainingstabilitydatacurriculumsyntheticreasoningannealinglongcontextextensionopenreproduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports YuLan-Mini, a 2.42B-parameter decoder-only base model pre-trained on 1.08T tokens of open and synthetic data. Its central claim is that a carefully engineered recipe — multi-stage data cleaning and a 27-phase curriculum, stability-focused initialization and optimization, and an 80B-token annealing stage with targeted selection and 28K context extension — lets this model match or beat industry small models trained on 2T to 18T tokens, including Qwen2.5-1.5B (18T), SmolLM2-1.7B (11T), and Llama3.2-3B (9T). If the claim holds, it means competitive base models can be reproduced at university scale, and the released per-phase data composition makes the recipe directly testable.

What carries the argument

The load-bearing mechanism is the three-part training recipe rather than any single technique. (1) The data pipeline combines MinHash de-duplication, heuristic filters, topic classifiers for math/code/reasoning recall, model-based quality scoring, n-gram decontamination, and large-scale synthetic generation of reasoning documents, chain-of-thought solutions, formal Lean proofs, and reflection data; these streams are arranged into 27 curriculum phases of 40B tokens each under a WSD schedule (10B warmup, 990B stable, 80B annealing) with per-phase mixture shifts kept under 3 percent. (2) The stability package scales initialization to $\sigma_{\mathrm{base}} = \sqrt{2/(5d)}$, scales the embedding output by 10, scales residual branches by $1.4\sqrt{n_{\mathrm{layers}}}$, applies µParameterization-style learning-rate scaling to QKV and FFN weights, and adds WeSaR reparameterization $W = \alpha \tilde{W}$ to decouple gradient size from direction, allowing a global learning rate of 0.01 with z-loss and a reduced Adam epsilon. (3) The annealing stage spends 80B tokens on a high-value mix selected by an accelerated gradient-based method (a LESS variant with InsTag), decays the learning rate with a 1-sqrt curve, and raises the RoPE base frequency from 10,000 to 490,000 to extend the context window to 28K tokens while using masked cross-document attention to preserve short-text performance.

What would settle it

Retrain YuLan-Mini from the released checkpoints and per-phase data under a pre-registered protocol — a single fixed chain-of-thought prompt, no curriculum feedback from the test sets, and baselines re-run in the same harness — and check whether the eight-benchmark average still beats Qwen2.5-1.5B and SmolLM2; if it falls behind when prompt choice and data-ratio tuning are taken away, the data-efficiency claim is not robust.

Watch

Extended reading notes

Core claim

YuLan-Mini is a 2.42B-parameter model with 56 layers, a 1,920-dimensional hidden width, grouped-query attention, and a 99K-vocabulary tokenizer with digit splitting. Trained on 1.08T tokens, the 28K-context checkpoint scores 37.80 on MATH-500 (4-shot), 64.00 on HumanEval (0-shot), 68.46 on GSM8K, and 49.10 on MMLU (5-shot). The authors argue that the model's aggregate performance, averaged over eight benchmarks, is competitive with small industry models trained on 2–18T tokens, despite using roughly one-half to one-seventeenth of their training budgets. The paper attributes this efficiency to a data pipeline that combines cleaning, classifier-based recall, decontamination, and a 27-phase WSD curriculum; to a stability package built on µParameterization-style initialization plus WeSaR reparameterization; and to an annealing stage that mixes high-value reasoning data with long-context training. It releases the full per-phase token composition to make the whole recipe reproducible.

Load-bearing premise

The entire data-efficiency comparison rests on the evaluation being a fair apples-to-apples contest: baseline scores come from each model's own paper, the better of two chain-of-thought prompts is picked per model, and the training-data mixture is adjusted based on the very benchmarks used in the final ranking, so if those numbers are not directly comparable the claimed token savings could shrink.

Editorial extensions

If this is right

  • A 2.42B base model can reach the top of its size class on math and code benchmarks after 1.08T tokens, with the 28K checkpoint scoring 37.80 on MATH-500, 68.46 on GSM8K, and 64.00 on HumanEval.
  • The 27-phase WSD curriculum, with per-phase mixture shifts capped at 3%, keeps training stable across 1.08T tokens at a global learning rate of 0.01.
  • The annealing stage extends the context window from 4K to 28K tokens by raising the RoPE base frequency to 490,000 while using long-context data with masked cross-document attention to preserve short-text skills.
  • Because the full per-phase data composition is released, the recipe can be reproduced and its components (data selection, annealing mix, stability package) can be ablated by other groups.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Whether the same 1.08T-token budget suffices at 7B scale is an open question; the stability package is designed to transfer via µParameterization, but the data curriculum and annealing are tuned for 2.4B.
  • The evaluation protocol's prompt selection (better of two per model) could inflate YuLan-Mini's margin; a single pre-registered prompt would make the data-efficiency claim sharper.
  • The annealing mix blends long-thought, formal-math, and gradient-selected data; isolating these components would show which one drives the MATH-500 and HumanEval gains.
  • The 4K and 28K checkpoints trade off general knowledge (MMLU 51.79 vs 49.10) for reasoning gains (MATH-500 32.60 vs 37.80); deployment choices may favor different checkpoints.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper reports the pre-training of YuLan-Mini, a 2.42B-parameter decoder-only base model trained on about 1.08T tokens, and claims that it reaches performance comparable to industry models trained on substantially more data. The technical recipe has three pillars: a data pipeline with cleaning, mixing, and curriculum scheduling; a stability-focused optimization setup combining scaled initialization, µP-like rules, and WeSaR re-parameterization; and an annealing stage with targeted data selection, learning-rate annealing, and context extension to 28K. The authors evaluate the model on a suite of math, code, commonsense, and Chinese benchmarks, and release detailed per-phase data compositions and checkpoints. The central claim is that the reported benchmark results demonstrate a data-efficient pre-training recipe reproducible in a university setting.

Significance. If the data-efficiency claim holds, the paper is a valuable resource for the community: it provides unusually detailed disclosure of data composition at each curriculum phase, a systematic discussion of training-stability diagnostics and mitigations, and an open release that lowers the barrier for reproducing competitive small models. The per-phase data tables and the stability analysis are concrete contributions that do not depend on the contested comparison. However, the headline claim depends on the evaluation protocol being a fair and controlled comparison, and that premise is weakened by the paper's own description of how the data mixture and annealing data were chosen.

major comments (5)
  1. [Section 4.5, Section 5.2, Tables 6-7] The central data-efficiency claim is not supported by a controlled evaluation protocol. Section 4.5 states that at each 40B-token curriculum boundary the data ratios are reassessed and adjusted based on the model's overall performance, with HumanEval given as an explicit example, and Section 5.2 states that formal math and o1-like reasoning data were incorporated to improve performance on challenging math benchmarks such as MATH-500. These are the same benchmarks used in Tables 6-7 and Figure 1 to argue for superior data efficiency. The reported math/code advantage may therefore reflect iterative tuning of the data mixture against the evaluation set rather than a generally more data-efficient pre-training recipe. I ask the authors to either hold out a set of benchmarks that were never used for data-mix decisions, report the trajectory of decisions together with the resulting scores, or evaluate frozen checkpoints from the originally scheduled curriculum.
  2. [Section 6.1.3, Tables 6-7] The comparison against baselines is not made under identical conditions. The text states that for CoT benchmarks the authors evaluate each model with both a short and a long prompt and select the higher score, while baseline numbers are mostly cited from official reports rather than re-run in the same harness (the table marks several values with an asterisk). This per-model prompt selection, combined with heterogeneous evaluation sources, can systematically inflate the relative standing of YuLan-Mini. The authors should re-evaluate all baselines with the identical evaluation code, generation limits, and a single fixed prompt per task, and report both prompt variants or justify why prompt selection cannot favor their model.
  3. [Section 5.1, Section 2.4] The annealing ratio and annealing function are fitted choices whose selection is part of the final model. The paper estimates the 8% annealing ratio from a scaling law and states that 1-sqrt annealing was chosen because it performed best empirically, and the final model is produced with that exact configuration. Since the headline comparison is made with the final model only, the reader cannot separate the effect of the annealing strategy from the effect of having selected the best-performing configuration on the evaluation benchmarks. At minimum, the paper should report the performance of checkpoints before annealing and, if available, results from alternative annealing ratios or functions.
  4. [Section 2.5 and Table 1] The paper's stability story relies partly on proxy-model experiments, but the transfer of the stability conclusions from the 0.05B/0.2B proxy models to the 2.42B model is asserted rather than demonstrated. Section 3.2.2 mentions that instability still appeared when migrating to the target size, and the mitigation is then validated mainly on small proxies and on the final successful run. Given that training-stability claims are one of the three advertised contributions, the authors should provide at least a controlled comparison showing that the chosen initialization/re-parameterization combination prevents divergence on the target-scale model under conditions where the baseline diverges.
  5. [Table 6 and Section 2.3] There is a discrepancy in the token count used for the main claim. The abstract and Section 2.3 state 1.08T tokens, but Table 6 reports the 4K checkpoint as trained on 1.04T tokens and the 28K checkpoint as trained on 1.08T tokens. Since the data-efficiency comparison in Figure 1 uses the average scores of the final model, the exact token count attributed to the evaluated checkpoint must be clarified. If the 4K checkpoint is the one used for parts of the comparison, the paper should state whether the comparison uses the 1.04T or 1.08T checkpoint.
minor comments (5)
  1. [Throughout] There are several typographical and grammatical errors: 'diffrent' in Table 1, 'intergration' in Section 2.5, 'have have' in Section 3.3.2, and 'hightlited' in Section 4.2. These should be corrected.
  2. [Section 2.4 and Table 3] The residual connection scaling factor is listed as 1.4√n_layers in Table 3 for YuLan-Mini, but the text in Section 3.2 does not derive or motivate this specific value; please add a short explanation or reference.
  3. [Figure 1] The caption says that models larger than 3B are plotted in gray, but the legend and axis labels do not make it easy to identify which points correspond to which models; please add labels to the points or a clearer legend.
  4. [Section 6.1.3] The evaluation section states that gpt-4o-mini is used to verify MATH-500 outputs and manual checks were conducted, but the number of samples checked manually is not given; please specify the verification procedure and sample size.
  5. [Appendix E] The detailed phase tables are useful, but the row for Phase 1 says the first 10B tokens are warmup and the next 30B are stable training; this should be stated directly above the table as well as in the main text for readability.

Circularity Check

2 steps flagged · score 6.0 of 10

Data-efficiency claim is partially circular: the data curriculum and annealing mix were tuned against the same benchmarks (HumanEval, MATH-500, GSM8K) that are then reported as evidence of superior data efficiency.

  1. fitted input called prediction [Section 4.5 (Data Curriculum), p.18; evidence used in Section 6.2 and Tables 6-7.]
    "For each 40B tokens, we reassess and adjust the data ratio when transitioning between training phases based on the model's overall performance in that phase. For example, if the model's performance on the HumanEval benchmark does not improve or declines after a stage, we may consider slightly increasing the amount of code data in the subsequent stage."

    The paper's headline claim of 'superior data efficiency' is supported by average scores on GSM8K, MATH-500, HumanEval, MBPP, MMLU, ARC-Challenge, HellaSwag, and CEval (Figure 1, Tables 6-7). Section 4.5 states that the data mixture itself was reassessed and adjusted every 40B tokens using the model's performance on benchmarks such as HumanEval. The final scores on those benchmarks are therefore partly optimized targets of the curriculum, not independent measurements of a fixed pretraining recipe. Reporting them as evidence that 1.08T tokens suffice is a fitted-input-called-prediction loop: the recipe was fit to the evaluation set and then evaluated on the same set.

  2. fitted input called prediction [Section 5.2 (Data Selection for Annealing Stage), p.19; compared with Table 6 MATH-500 row.]
    "In particular, we incorporate formal mathematical reasoning (theorem proving in Lean) and advanced reasoning data (o1-like thought data) to improve the model's performance on challenging math benchmarks, e.g., MATH-500, which have been shown in Table 6."

    This sentence explicitly states that annealing-stage data were selected with the goal of improving MATH-500, and then points to Table 6, where YuLan-Mini's MATH-500 score (37.80) is reported as evidence of mathematical capability and of training efficacy. The MATH-500 gain is the objective of the data-selection procedure, so citing it as an independent confirmation of the annealing approach is circular: the benchmark improvement is what the selection was engineered to produce, not a predicted consequence of a general recipe.

full rationale

The paper is transparent and mostly self-contained about architecture, data, and optimization choices; the training-stability analysis in Section 3 is an independent empirical study with proxy models, and the annealing-ratio estimate (8%) is taken from an external scaling law (Tissue et al., 2024), which is not circular. However, the central data-efficiency claim is weakened by a genuine feedback loop that the paper itself documents. Section 4.5 says data ratios were reassessed and adjusted at each 40B boundary based on model performance on benchmarks, with HumanEval as the example; Section 5.2 says annealing data were chosen to improve MATH-500. The same benchmarks are then used in Tables 6-7 and Figure 1 to argue for 'superior training efficacy' from only 1.08T tokens. Consequently, the reported math/code advantages are at least partially optimized targets rather than independent evidence of a generally data-efficient recipe. The evaluation section also acknowledges a second limitation: baseline scores are cited from official reports rather than re-run under the identical protocol ('fully reproducing the results of these baseline models as originally reported remains challenging'), so the comparison does not control for this tuning loop. This is an experimental-design confound, not an allegation of misconduct, and it does not affect the paper's other contributions such as the stability methods, data release, and curriculum transparency. Because two of the headline 'predictions' (data efficiency on math/code benchmarks) reduce to benchmark-driven fitting, a score of 6 is appropriate.

Assumptions & free parameters 8 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a set of engineering choices (learning rate, annealing ratio, data mix, scaling factors) that are fitted or hand-picked, and on domain assumptions about the validity of cited baselines and synthetic data. No new physical or mathematical entities are introduced.

free parameters (8)
  • Learning rate = 0.01
    Selected via proxy model search (Section 2.4, Appendix B).
  • Warmup tokens = 10B
    Set based on Wortsman et al. (2024) recommendation to extend warmup; Section 2.4.
  • Annealing ratio = 8% (80B tokens)
    Estimated using the scaling law of learning rate annealing from Tissue et al. (2024); Section 2.4.
  • RoPE base frequency (annealing) = 490,000
    Chosen to extend context to 28K tokens; Section 5.1.
  • Embedding scaling factor = 10
    Empirically chosen to keep LN input variance near 1; Section 3.1.2.
  • Data mixture proportions = 60% English, 20% code, 10% math, 10% Chinese, with per-phase adjustments up to 3%
    Adjusted during training based on benchmark performance and validation PPL; Section 4.4 and 4.5.
  • BPE dropout rate = 0.2
    Chosen as a 'relatively low' rate; Section 2.2.
  • Z-loss coefficient = 1e-4
    Set to encourage logits near zero; Section 3.3.2.
assumptions (6)
  • standard math Transformer architecture and next-token prediction objective are effective for language modeling.
    Implicit in Sections 1-2; no derivation provided.
  • domain assumption Scaling laws (Kaplan et al.) accurately estimate FLOPs and annealing behavior.
    Used in Figure 1 and for setting annealing ratio (Section 2.4).
  • domain assumption Benchmark scores from different papers are comparable when evaluated under similar settings.
    Baseline numbers cited from official reports; Section 6.1.2.
  • domain assumption Synthetic data from Qwen2.5-Math and QwQ-32B-Preview does not introduce benchmark contamination beyond the n-gram decontamination.
    Used extensively in Section 4.3; decontamination in Section 4.2.
  • ad hoc to paper Hidden states variance and gradient norm are reliable indicators of training instability that transfer from 0.2B proxy models to 2.42B model.
    Section 3.1; no external verification that this indicator is sufficient.
  • ad hoc to paper The 1-sqrt annealing function outperforms alternatives.
    Empirically claimed in Section 5.1 without quantitative comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of YuLan-Mini: An Open Data-efficient Language Model." pith.science (2026). https://pith.science/paper/H2ZKK26I

@misc{pith2026241217743,
  author       = {Pith},
  title        = {Pith review of: YuLan-Mini: An Open Data-efficient Language Model},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H2ZKK26I}},
  note         = {Machine review of arXiv:2412.17743}
}
read the original abstract

Effective pre-training of large language models (LLMs) has been challenging due to the immense resource demands and the complexity of the technical processes involved. This paper presents a detailed technical report on YuLan-Mini, a highly capable base model with 2.42B parameters that achieves top-tier performance among models of similar parameter scale. Our pre-training approach focuses on enhancing training efficacy through three key technical contributions: an elaborate data pipeline combines data cleaning with data schedule strategies, a robust optimization method to mitigate training instability, and an effective annealing approach that incorporates targeted data selection and long context training. Remarkably, YuLan-Mini, trained on 1.08T tokens, achieves performance comparable to industry-leading models that require significantly more data. To facilitate reproduction, we release the full details of the data composition for each training phase. Project details can be accessed at the following link: https://github.com/RUC-GSAI/YuLan-Mini.

Figures

Figures reproduced from arXiv: 2412.17743 by the authors.

Figure 1
Figure 1. Performance comparison of YuLan-Mini against other [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Training loss and gradients during pre-training process. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Comparison of training dynamics between divergent and convergent trial. The [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Variance of LN output of each layers. 0 2000 4000 6000 8000 Training steps 10 1 10 2 10 3 10 4 10 5 Mean of attention scores Mean of attention scores Variance of LayerNorm output 0.1 1 Variance of LayerNorm output [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 6
Figure 6. Figure 6: Ablation experiments on training instability mitigation methods are conducted. We report [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: The curves of attention value and LN output variances (left) and gradient norm and loss [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Illustration of our data filtering pipeline and synthetic generation for reasoning data. The [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: The data mixture proportion of math, code, and general data. We keep the proportion of web data unchanged in stable stage, and then gradually decrease it in annealing stage. The entire process is divided into into three major stages: warmup, stable training, and anneal…
Figure 1
Figure 1. Figure 1: Based on these results, we can identify the following key observations: [PITH_FULL_IMAGE:figures/full_fig_p023_1.png]
Figure 10
Figure 10. Figure 10: Performance comparison using perplexity (PPL) and accuracy-based metrics to monitor [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Falcon-H1 reports competitive benchmark scores for a 0.5B to 34B family of parallel hybrid attention/Mamba-2 models, claiming 2x to 4x parameter efficiency versus dense transformers.

Reference graph

Works this paper leans on

135 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Synthetically generated reasoning dataset (gsm8k-inspired) with enhanced diversity using gretel navigator and meta-llama/meta-llama-3.1-405b

    Gretel AI. Synthetically generated reasoning dataset (gsm8k-inspired) with enhanced diversity using gretel navigator and meta-llama/meta-llama-3.1-405b. https://huggingface.co/gretelai/synthetic-gsm8k-reflection-405b, 9 2024

  2. [2]

    GQA: training generalized multi-query transformer models from multi-head checkpoints

    Joshua Ainslie, James Lee - Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebr \' o n, and Sumit Sanghai. GQA: training generalized multi-query transformer models from multi-head checkpoints. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singa...

  3. [3]

    Smollm2 - with great data, comes great performance, 2024

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Lewis Tunstall, Agustín Piqueres, Andres Marafioti, Cyril Zakka, Leandro von Werra, and Thomas Wolf. Smollm2 - with great data, comes great performance, 2024

  4. [4]

    Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J

    Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. Program synthesis with large language models. CoRR, abs/2108.07732, 2021. URL https://arxiv.org/abs/2108.07732

  5. [5]

    Jiang, Jia Deng, Stella Biderman, and Sean Welleck

    Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen Marcus McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://open...

  6. [6]

    Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer normalization. CoRR, abs/1607.06450, 2016. URL http://arxiv.org/abs/1607.06450

  7. [7]

    Numinamath 7b cot, 2024

    Edward Beeching, Shengyi Costa Huang, Albert Jiang, Jia Li, Benjamin Lipkin, Zihan Qina, Kashif Rasul, Ziju Shen, Roman Soletskyi, and Lewis Tunstall. Numinamath 7b cot, 2024. URL http://faculty.bicmr.pku.edu.cn/ dongbin/Publications/numina_dataset.pdf

  8. [8]

    Stable LM 2 1.6b technical report

    Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparaju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccol \' o Zanichelli, and Carlos Riquelme. Stable LM 2 1.6b technical report. CoRR, abs/2402.17834...

Show all 135 references
  1. [9]

    Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, Guangbo Hao, Zhewen Hao, Ying He, Wenjie Hu, Panpan Huang, Erhang Li, Guowei ...

  2. [10]

    Enriching word vectors with subword information

    Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5: 0 135--146, 2017. ISSN 2307-387X

  3. [11]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  4. [12]

    Towards effective and efficient continual pre-training of large language models

    Jie Chen, Zhipeng Chen, Jiapeng Wang, Kun Zhou, Yutao Zhu, Jinhao Jiang, Yingqian Min, Wayne Xin Zhao, Zhicheng Dou, Jiaxin Mao, Yankai Lin, Ruihua Song, Jun Xu, Xu Chen, Rui Yan, Zhewei Wei, Di Hu, Wenbing Huang, and Ji - Rong Wen. Towards effective and efficient continual pr...

  5. [13]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Sco...

  6. [14]

    Extending context window of large language models via positional interpolation

    Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. Extending context window of large language models via positional interpolation. CoRR, abs/2306.15595, 2023. doi:10.48550/ARXIV.2306.15595. URL https://doi.org/10.48550/arXiv.2306.15595

  7. [15]

    Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vino...

  8. [16]

    Stable language model pre-training by reducing embedding variability

    Woojin Chung, Jiwoo Hong, Na Min An, James Thorne, and Se - Young Yun. Stable language model pre-training by reducing embedding variability. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural La...

  9. [17]

    Training verifiers to solve math word problems

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://...

  10. [18]

    Getting the most out of your tokenizer for pre-training and domain adaptation

    Gautier Dagan, Gabriel Synnaeve, and Baptiste Rozi \` e re. Getting the most out of your tokenizer for pre-training and domain adaptation. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL http...

  11. [19]

    Flashattention-2: Faster attention with better parallelism and work partitioning

    Tri Dao. Flashattention-2: Faster attention with better parallelism and work partitioning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL https://openreview.net/forum?id=mZn2Xyh9Ec

  12. [20]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e . Flashattention: Fast and memory-efficient exact attention with io-awareness. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Proces...

  13. [21]

    The z-loss: a shift and scale invariant classification loss belonging to the spherical family

    Alexandre de Br \' e bisson and Pascal Vincent. The z-loss: a shift and scale invariant classification loss belonging to the spherical family. CoRR, abs/1604.08859, 2016. URL http://arxiv.org/abs/1604.08859

  14. [22]

    DeepSeek - AI, Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Deng, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, Hao Zhang, Hanwei Xu, Hao Yang, Haow...

  15. [23]

    Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster

    Nolan Dey, Gurpreet Gosal, Zhiming Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness. Cerebras-gpt: Open compute-optimal language models trained on the cerebras wafer-scale cluster. CoRR, abs/2304.03208, 2023 a . doi:10.48550/ARXIV.2304.0320...

  16. [24]

    Cerebras- GPT : Open Compute - Optimal Language Models Trained on the Cerebras Wafer - Scale Cluster , April 2023 b

    Nolan Dey, Gurpreet Gosal, Zhiming, Chen, Hemant Khachane, William Marshall, Ribhu Pathria, Marvin Tom, and Joel Hestness. Cerebras- GPT : Open Compute - Optimal Language Models Trained on the Cerebras Wafer - Scale Cluster , April 2023 b . URL http://arxiv.org/abs/2304.03208....

  17. [25]

    Fewer truncations improve language modeling

    Hantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar, Anoop Deoras, Dan Roth, and Stefano Soatto. Fewer truncations improve language modeling. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024 a...

  18. [26]

    Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , October 2024 b

    Yuyang Ding, Xinyu Shi, Xiaobo Liang, Juntao Li, Qiaoming Zhu, and Min Zhang. Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch , October 2024 b . URL http://arxiv.org/abs/2410.18693. arXiv:2410.18693 [cs]

  19. [27]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al - Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Z...

  20. [28]

    How to train long-context language models (effectively)

    Tianyu Gao, Alexander Wettig, Howard Yen, and Danqi Chen. How to train long-context language models (effectively). CoRR, abs/2410.02660, 2024. doi:10.48550/ARXIV.2410.02660. URL https://doi.org/10.48550/arXiv.2410.02660

  21. [29]

    Dirk Groeneveld, Iz Beltagy, Evan Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu,...

  22. [30]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. Deepseek-coder: When the large language model meets programming - the rise of code intelligence. CoRR, abs/2401.14196, 202...

  23. [31]

    Scaling laws and compute-optimal training beyond fixed training durations

    Alexander H \" a gele, Elie Bakouch, Atli Kosson, Loubna Ben Allal, Leandro von Werra, and Martin Jaggi. Scaling laws and compute-optimal training beyond fixed training durations. CoRR, abs/2405.18392, 2024. doi:10.48550/ARXIV.2405.18392. URL https://doi.org/10.48550/arXiv.2405.18392

  24. [32]

    Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning

    Xiaotian Han, Yiren Jian, Xuefeng Hu, Haogeng Liu, Yiqi Wang, Qihang Fan, Yuang Ai, Huaibo Huang, Ran He, Zhenheng Yang, and Quanzeng You. Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning. CoRR, abs/2409.12568, 2024. doi:10.48550/ARXIV....

  25. [33]

    Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models

    Conghui He, Zhenjiang Jin, Chao Xu, Jiantao Qiu, Bin Wang, Wei Li, Hang Yan, Jiaqi Wang, and Dahua Lin. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models. CoRR, abs/2308.10755, 2023. doi:10.48550/ARXIV.2308.10755. URL https://doi.org/10...

  26. [34]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview...

  27. [35]

    Measuring mathematical problem solving with the MATH dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Joaquin Vanschoren and Sai - Kit Yeung, editors, Proceedings of the Neural Information Processi...

  28. [36]

    RULER: what's the real context size of your long-context language models? CoRR, abs/2404.06654, 2024

    Cheng - Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. RULER: what's the real context size of your long-context language models? CoRR, abs/2404.06654, 2024. doi:10.48550/ARXIV.2404.06654. URL https://doi.org/10.48...

  29. [37]

    Liger kernel: Efficient triton kernels for LLM training

    Pin - Lun Hsu, Yun Dai, Vignesh Kothapalli, Qingquan Song, Shao Tang, Siyu Zhu, Steven Shimizu, Shivam Sahni, Haowen Ning, and Yanning Chen. Liger kernel: Efficient triton kernels for LLM training. CoRR, abs/2410.10989, 2024. doi:10.48550/ARXIV.2410.10989. URL https://doi.org/...

  30. [38]

    Minicpm: Unveiling the potential of small language models with scalable training strategies

    Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, Xinrong Zhang, Zhen Leng Thai, Kai Zhang, Chongyi Wang, Yuan Yao, Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai Li, Z...

  31. [39]

    Towards reasoning in large language models: A survey

    Jie Huang and Kevin Chen - Chuan Chang. Towards reasoning in large language models: A survey. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 104...

  32. [40]

    Siming Huang, Tianhao Cheng, J. K. Liu, Jiaran Hao, Liuyihan Song, Yang Xu, J. Yang, J. H. Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Zhaoxiang Zhang, Jie Fu, Qian Liu, Ge Zhang, Zili Wang, Yuan Qi, Yinghui Xu, and Wei Chu. OpenCoder : The Open Cookbook for Top - Tier C...

  33. [41]

    C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

    Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models. In Advances in Neural ...

  34. [42]

    Technical report: Enhancing llm reasoning with reward-guided tree search

    Jinhao Jiang, Zhipeng Chen, Yingqian Min, Jie Chen, Xiaoxue Cheng, Jiapeng Wang, Yiru Tang, Haoxiang Sun, Jia Deng, Wayne Xin Zhao, Zheng Liu, Dong Yan, Jian Xie, Zhongyuan Wang, and Ji-Rong Wen. Technical report: Enhancing llm reasoning with reward-guided tree search. CoRR, a...

  35. [43]

    TinyBERT : Distilling BERT for Natural Language Understanding , October 2020

    Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. TinyBERT : Distilling BERT for Natural Language Understanding , October 2020. URL http://arxiv.org/abs/1909.10351. Issue: arXiv:1909.10351 1097 citations (Semantic Scholar/arXiv) [2...

  36. [44]

    Calc-x and calcformers: Empowering arithmetical chain-of-thought through interaction with symbolic systems

    Marek Kadlc \' k, Michal Stef \' a nik, Ondrej Sotol \' a r, and Vlastimil Martinek. Calc-x and calcformers: Empowering arithmetical chain-of-thought through interaction with symbolic systems. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Confe...

  37. [45]

    Kaplan, Sam McCandlish, T

    J. Kaplan, Sam McCandlish, T. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. Scaling Laws for Neural Language Models . ArXiv, January 2020. URL https://www.semanticscholar.org/paper/Scaling-Laws-for-Neural-Language-Mod...

  38. [46]

    LAMBADA: backward chaining for automated reasoning in natural language

    Mehran Kazemi, Najoung Kim, Deepti Bhatia, Xin Xu, and Deepak Ramachandran. LAMBADA: backward chaining for automated reasoning in natural language. In Anna Rogers, Jordan L. Boyd - Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association f...

  39. [47]

    Strategic data ordering: Enhancing large language model performance through curriculum learning

    Jisu Kim and Juhwan Lee. Strategic data ordering: Enhancing large language model performance through curriculum learning. CoRR, abs/2405.07490, 2024. doi:10.48550/ARXIV.2405.07490. URL https://doi.org/10.48550/arXiv.2405.07490

  40. [48]

    Efficient memory management for large language model serving with pagedattention

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Jason Flinn, Margo I. Seltzer, Peter Druschel, Antoine Kaufmann, and...

  41. [49]

    Race: Large-scale reading comprehension dataset from examinations, 2017

    Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations, 2017. URL https://arxiv.org/abs/1704.04683

  42. [50]

    Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...

  43. [51]

    To FP8 and back again: Quantifying the effects of reducing precision on LLM training stability

    Joonhyung Lee, Jeongin Bae, Byeongwook Kim, Se Jung Kwon, and Dongsoo Lee. To FP8 and back again: Quantifying the effects of reducing precision on LLM training stability. CoRR, abs/2405.18710, 2024. doi:10.48550/ARXIV.2405.18710. URL https://doi.org/10.48550/arXiv.2405.18710

  44. [52]

    The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models

    Conglong Li, Minjia Zhang, and Yuxiong He. The stability-efficiency dilemma: Investigating sequence length warmup for training GPT models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems ...

  45. [53]

    CMMLU: measuring massive multitask language understanding in chinese

    Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. CMMLU: measuring massive multitask language understanding in chinese. In Lun - Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computationa...

  46. [54]

    Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Raghavi Chandu, Thao Nguyen, Igor Vasiljevic, Sham M

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Alb...

  47. [55]

    Numinamath

    Jia LI, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Costa Huang, Kashif Rasul, Longhui Yu, Albert Jiang, Ziju Shen, Zihan Qin, Bin Dong, Li Zhou, Yann Fleureau, Guillaume Lample, and Stanislas Polu. Numinamath. [https://github.com/project-numina/aimo-...

  48. [56]

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy - Poirier, Jo \ a o Mont...

  49. [57]

    Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023

    Wing Lian, Guan Wang, Bleys Goodson, Eugene Pentland, Austin Cook, Chanvichet Vong, and "Teknium". Slimorca: An open dataset of gpt-4 augmented flan reasoning traces, with verification, 2023. URL https://https://huggingface.co/Open-Orca/SlimOrca

  50. [58]

    Universal checkpointing: Efficient and flexible checkpointing for large scale distributed training

    Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko, Masahiro Tanaka, Stas Bekman, Olatunji Ruwase, and Minjia Zhang. Universal checkpointing: Efficient and flexible checkpointing for large scale distributed training. CoRR, abs/2406.18820, 2024. doi:10.48550/ARXIV.2406.18820. URL https:...

  51. [59]

    Let's verify step by step

    Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let's verify step by step. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-1...

  52. [60]

    Lean-star: Learning to interleave thinking and proving

    Haohan Lin, Zhiqing Sun, Yiming Yang, and Sean Welleck. Lean-star: Learning to interleave thinking and proving. CoRR, abs/2407.10040, 2024. doi:10.48550/ARXIV.2407.10040. URL https://doi.org/10.48550/arXiv.2407.10040

  53. [61]

    Evaluating language models for efficient code generation

    Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. Evaluating language models for efficient code generation. In First Conference on Language Modeling, 2024 a . URL https://openreview.net/forum?id=IBCBMeAhmC

  54. [62]

    Longwanjuan: Towards systematic measurement for long text quality

    Xiaoran Liu, Kai Lv, Qipeng Guo, Hang Yan, Conghui He, Xipeng Qiu, and Dahua Lin. Longwanjuan: Towards systematic measurement for long text quality. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Findings of the Association for Computational Linguistics: EM...

  55. [63]

    Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra

    Zechun Liu, Changsheng Zhao, Forrest N. Iandola, Chen Lai, Yuandong Tian, Igor Fedorov, Yunyang Xiong, Ernie Chang, Yangyang Shi, Raghuraman Krishnamoorthi, Liangzhen Lai, and Vikas Chandra. Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. I...

  56. [64]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openreview.net/forum?id=Bkg6RiCqY7

  57. [65]

    Fineweb-edu, May 2024 a

    Anton Lozhkov, Loubna Ben Allal, Leandro von Werra, and Thomas Wolf. Fineweb-edu, May 2024 a . URL https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu

  58. [66]

    McAuley, Han Hu, Torsten Scholak, S \' e bastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, and et al

    Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy - Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul,...

  59. [67]

    \# instag: Instruction tagging for analyzing supervised fine-tuning of large language models

    Keming Lu, Hongyi Yuan, Zheng Yuan, Runji Lin, Junyang Lin, Chuanqi Tan, Chang Zhou, and Jingren Zhou. \# instag: Instruction tagging for analyzing supervised fine-tuning of large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, ...

  60. [68]

    Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivi \` e re, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, L \' e onard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro - Ros, Ambrose ...

  61. [69]

    Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems

    Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, Wayne Xin Zhao, Zheng Liu, Zhongyuan Wang, and Ji-Rong Wen. Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning system...

  62. [70]

    Agentinstruct: Toward generative teaching with agentic flows

    Arindam Mitra, Luciano Del Corro, Guoqing Zheng, Shweti Mahajan, Dany Rouhana, Andr \' e s Codas, Yadong Lu, Weige Chen, Olga Vrousgos, Corby Rosset, Fillipe Silva, Hamed Khanpour, Yash Lara, and Ahmed Awadallah. Agentinstruct: Toward generative teaching with agentic flows. Co...

  63. [71]

    Orca-math: Unlocking the potential of slms in grade school math

    Arindam Mitra, Hamed Khanpour, Corby Rosset, and Ahmed Awadallah. Orca-math: Unlocking the potential of slms in grade school math. CoRR, abs/2402.14830, 2024 b . doi:10.48550/ARXIV.2402.14830. URL https://doi.org/10.48550/arXiv.2402.14830

  64. [72]

    A theory on adam instability in large-scale machine learning

    Igor Molybog, Peter Albert, Moya Chen, Zachary DeVito, David Esiobu, Naman Goyal, Punit Singh Koura, Sharan Narang, Andrew Poulton, Ruan Silva, Binh Tang, Diana Liskovich, Puxin Xu, Yuchen Zhang, Melanie Kambadur, Stephen Roller, and Susan Zhang. A theory on adam instability i...

  65. [73]

    A corpus and evaluation framework for deeper understanding of commonsense stories, 2016

    Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. A corpus and evaluation framework for deeper understanding of commonsense stories, 2016. URL https://arxiv.org/abs/1604.01696

  66. [74]

    Initialization of large language models via reparameterization to mitigate loss spikes

    Kosuke Nishida, Kyosuke Nishida, and Kuniko Saito. Initialization of large language models via reparameterization to mitigate loss spikes. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Lang...

  67. [75]

    GPT-4 technical report

    OpenAI. GPT-4 technical report. CoRR, abs/2303.08774, 2023. doi:10.48550/ARXIV.2303.08774. URL https://doi.org/10.48550/arXiv.2303.08774

  68. [76]

    Opencsg/chinese-fineweb-edu Datasets at Hugging Face

    Opencsg. Opencsg/chinese-fineweb-edu Datasets at Hugging Face . https://huggingface.co/datasets/opencsg/chinese-fineweb-edu

  69. [77]

    Openwebmath: An open dataset of high-quality mathematical web text

    Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. Openwebmath: An open dataset of high-quality mathematical web text. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024. URL htt...

  70. [78]

    The FineWeb Datasets : Decanting the Web for the Finest Text Data at Scale , October 2024

    Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The FineWeb Datasets : Decanting the Web for the Finest Text Data at Scale , October 2024. URL http://arxiv.org/abs/2406.17557. arXiv:2406.17557

  71. [79]

    Using the output embedding to improve language models

    Ofir Press and Lior Wolf. Using the output embedding to improve language models. In Mirella Lapata, Phil Blunsom, and Alexander Koller, editors, Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2017, Valencia, Sp...

  72. [80]

    Bpe-dropout: Simple and effective subword regularization

    Ivan Provilkov, Dmitrii Emelianenko, and Elena Voita. Bpe-dropout: Simple and effective subword regularization. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistic...

  73. [81]

    Qwen2.5: A party of foundation models, September 2024

    Qwen-Team . Qwen2.5: A party of foundation models, September 2024. URL https://qwenlm.github.io/blog/qwen2.5/

  74. [82]

    Qwq: Reflect deeply on the boundaries of the unknown, November 2024

    Qwen-Team. Qwq: Reflect deeply on the boundaries of the unknown, November 2024. URL https://qwenlm.github.io/blog/qwq-32b-preview/

  75. [83]

    Zero: memory optimizations toward training trillion parameter models

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. Zero: memory optimizations toward training trillion parameter models. In Christine Cuicchi, Irene Qualters, and William T. Kramer, editors, Proceedings of the International Conference for High Performance Comput...

  76. [84]

    Methods of improving LLM training stability

    Oleg Rybakov, Mike Chrzanowski, Peter Dykas, Jinze Xue, and Ben Lanir. Methods of improving LLM training stability. CoRR, abs/2410.16682, 2024. doi:10.48550/ARXIV.2410.16682. URL https://doi.org/10.48550/arXiv.2410.16682

  77. [85]

    Winogrande: an adversarial winograd schema challenge at scale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: an adversarial winograd schema challenge at scale. Commun. ACM , 64 0 (9): 0 99--106, 2021. doi:10.1145/3474381. URL https://doi.org/10.1145/3474381

  78. [86]

    Analysing mathematical reasoning abilities of neural models

    David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. Analysing mathematical reasoning abilities of neural models. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. URL https://openr...

  79. [87]

    Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M. Saiful Bari, Stella Biderman, Hady Elsahar, Niklas Muennighoff, Jason Phang, Ofir Press, Colin Raffel, Victor Sanh, Sheng Shen, Lintang Sutawika, Jaesung Tae, Zheng Xin Yong, Julien Launay, and Iz Bel...

  80. [88]

    GLU variants improve transformer

    Noam Shazeer. GLU variants improve transformer. CoRR, abs/2002.05202, 2020. URL https://arxiv.org/abs/2002.05202

  81. [89]

    Reflexion: language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: language agents with verbal reinforcement learning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information...

  82. [90]

    Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020. URL https://arxiv.org/abs/1909.08053

  83. [91]

    Scaling synthetic logical reasoning datasets with context-sensitive declarative grammars

    Damien Sileo. Scaling synthetic logical reasoning datasets with context-sensitive declarative grammars. In Yaser Al - Onaizan, Mohit Bansal, and Yun - Nung Chen, editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami,...

  84. [92]

    Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, N...

  85. [93]

    An integrated data processing framework for pretraining foundation models

    Yiding Sun, Feng Wang, Yutao Zhu, Wayne Xin Zhao, and Jiaxin Mao. An integrated data processing framework for pretraining foundation models. In Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang, editors, Proceedings of the 47th International ACM...

  86. [94]

    Spike no more: Stabilizing the pre-training of large language models

    Sho Takase, Shun Kiyono, Sosuke Kobayashi, and Jun Suzuki. Spike no more: Stabilizing the pre-training of large language models. CoRR, abs/2312.16903, 2023. doi:10.48550/ARXIV.2312.16903. URL https://doi.org/10.48550/arXiv.2312.16903

  87. [95]

    Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...

  88. [96]

    LLMBox : A Comprehensive Library for Large Language Models

    Tianyi Tang, Hu Yiwen, Bingqian Li, Wenyang Luo, ZiJing Qin, Haoxiang Sun, Jiapeng Wang, Shiyi Xu, Xiaoxue Cheng, Geyang Guo, Han Peng, Bowen Zheng, Yiru Tang, Yingqian Min, Yushuo Chen, Jie Chen, Ranchi Zhao, Luran Ding, Yuhao Wang, Zican Dong, Xia Chunxuan, Junyi Li, Kun Zho...

  89. [97]

    Mathscale: Scaling instruction tuning for mathematical reasoning

    Zhengyang Tang, Xingxing Zhang, Benyou Wang, and Furu Wei. Mathscale: Scaling instruction tuning for mathematical reasoning. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024 c . URL https://openrev...

  90. [98]

    Gemma Team. Gemma. 2024. doi:10.34740/KAGGLE/M/3301. URL https://www.kaggle.com/m/3301

  91. [99]

    D4: improving LLM pretraining via document de-duplication and diversification

    Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari Morcos. D4: improving LLM pretraining via document de-duplication and diversification. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information P...

  92. [100]

    Scaling law with learning rate annealing

    Howe Tissue, Venus Wang, and Lu Wang. Scaling law with learning rate annealing. CoRR, abs/2408.11029, 2024. doi:10.48550/ARXIV.2408.11029. URL https://doi.org/10.48550/arXiv.2408.11029

  93. [101]

    Openmathinstruct-1: A 1.8 million math instruction tuning dataset

    Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, and Igor Gitman. Openmathinstruct-1: A 1.8 million math instruction tuning dataset. CoRR, abs/2402.10176, 2024. doi:10.48550/ARXIV.2402.10176. URL https://doi.org/10.48550/arXiv.2402.10176

  94. [102]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie - Anne Lachaux, Timoth \' e e Lacroix, Baptiste Rozi \` e re, Naman Goyal, Eric Hambro, Faisal Azhar, Aur \' e lien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient fo...

  95. [103]

    Tokenization matters! degrading large language models through challenging their tokenization

    Dixuan Wang, Yanda Li, Junyuan Jiang, Zepeng Ding, Guochao Jiang, Jiaqing Liang, and Deqing Yang. Tokenization matters! degrading large language models through challenging their tokenization. CoRR, abs/2405.17067, 2024 a . doi:10.48550/ARXIV.2405.17067. URL https://doi.org/10....

  96. [104]

    Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning

    Ke Wang, Houxing Ren, Aojun Zhou, Zimu Lu, Sichun Luo, Weikang Shi, Renrui Zhang, Linqi Song, Mingjie Zhan, and Hongsheng Li. Mathcoder: Seamless code integration in llms for enhanced mathematical reasoning. In The Twelfth International Conference on Learning Representations, ...

  97. [105]

    How do your code llms perform? empowering code instruction tuning with really good data

    Yejie Wang, Keqing He, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, and Weiran Xu. How do your code llms perform? empowering code instruction tuning with really good data. In Yaser A...

  98. [106]

    Chi, Quoc V

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. In Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh, editors, ...

  99. [107]

    Magicoder: Empowering code generation with oss-instruct

    Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. Magicoder: Empowering code generation with oss-instruct. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. URL https://openreview...

  100. [108]

    Liu, Lechao Xiao, Katie E

    Mitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett, Alexander A. Alemi, Ben Adlam, John D. Co - Reyes, Izzeddin Gur, Abhishek Kumar, Roman Novak, Jeffrey Pennington, Jascha Sohl - Dickstein, Kelvin Xu, Jaehoon Lee, Justin Gilmer, and Simon Kornblith. Small-scale pr...

  101. [109]

    Rabe, Wenda Li, Jimmy Ba, Roger B

    Yuhuai Wu, Markus N. Rabe, Wenda Li, Jimmy Ba, Roger B. Grosse, and Christian Szegedy. LIME: learning inductive bias for primitives of mathematical reasoning. In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 20...

  102. [110]

    Lean-github: Compiling github LEAN repositories for a versatile LEAN prover

    Zijian Wu, Jiayu Wang, Dahua Lin, and Kai Chen. Lean-github: Compiling github LEAN repositories for a versatile LEAN prover. CoRR, abs/2407.17227, 2024. doi:10.48550/ARXIV.2407.17227. URL https://doi.org/10.48550/arXiv.2407.17227

  103. [111]

    LESS: selecting influential data for targeted instruction tuning

    Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. LESS: selecting influential data for targeted instruction tuning. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 . OpenReview.net, 2024. ...

  104. [112]

    Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data

    Huajian Xin, Daya Guo, Zhihong Shao, Zhizhou Ren, Qihao Zhu, Bo Liu, Chong Ruan, Wenda Li, and Xiaodan Liang. Deepseek-prover: Advancing theorem proving in llms through large-scale synthetic data. CoRR, abs/2405.14333, 2024. doi:10.48550/ARXIV.2405.14333. URL https://doi.org/1...

  105. [113]

    On layer normalization in the transformer architecture

    Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie - Yan Liu. On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 ...

  106. [114]

    Effective long-context scaling of foundation models

    Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, S...

  107. [115]

    Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing

    Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing. CoRR, abs/2406.08464, 2024. doi:10.48550/ARXIV.2406.08464. URL https://doi.org/10.485...

  108. [116]

    Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering

    Vikas Yadav, Steven Bethard, and Mihai Surdeanu. Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors, Proceedings of the 2019 Conference on Empirical Met...

  109. [117]

    Baichuan 2: Open large-scale language models

    Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Ce Bian, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, Fei Deng, Feng Wang, Feng Liu, Guangwei Ai, Guosheng Dong, Haizhou Zhao, Hang Xu, Haoze Sun, Hongda Zhang, Hui Liu, Jiaming Ji, Jian Xie, Juntao Dai, Kun Fa...

  110. [118]

    Qwen2 technical report

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...

  111. [119]

    Qwen2.5-math technical report: Toward mathematical expert model via self-improvement

    An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Keming Lu, Mingfeng Xue, Runji Lin, Tianyu Liu, Xingzhang Ren, and Zhenru Zhang. Qwen2.5-math technical report: Toward mathematical expert model via se...

  112. [120]

    Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao

    Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tensor programs V: tuning large neural networks via zero-shot hyperparameter transfer. CoRR, abs/2203.03466, 2022. doi:10.48550/ARXIV.2...

  113. [121]

    Tensor programs VI: feature learning in infinite depth neural networks

    Greg Yang, Dingli Yu, Chen Zhu, and Soufiane Hayou. Tensor programs VI: feature learning in infinite depth neural networks. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview.net, 2024 c . URL https://op...

  114. [122]

    Lean workbook: A large-scale lean problem set formalized from natural language math problems

    Huaiyuan Ying, Zijian Wu, Yihan Geng, Jiayu Wang, Dahua Lin, and Kai Chen. Lean workbook: A large-scale lean problem set formalized from natural language math problems. CoRR, abs/2406.03847, 2024. doi:10.48550/ARXIV.2406.03847. URL https://doi.org/10.48550/arXiv.2406.03847

  115. [123]

    Yoo, Morris A

    Andy B. Yoo, Morris A. Jette, and Mark Grondona. SLURM: simple linux utility for resource management. In Dror G. Feitelson, Larry Rudolph, and Uwe Schwiegelshohn, editors, Job Scheduling Strategies for Parallel Processing, 9th International Workshop, JSSPP 2003, Seattle, WA, U...

  116. [124]

    Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

    Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representation...

  117. [125]

    Mammoth: Building math generalist models through hybrid instruction tuning

    Xiang Yue, Xingwei Qu, Ge Zhang, Yao Fu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mammoth: Building math generalist models through hybrid instruction tuning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024 ....

  118. [126]

    Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Llu \' s M \` a rquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguisti...

  119. [127]

    Root mean square layer normalization

    Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d'Alch \' e - Buc, Emily B. Fox, and Roman Garnett, editors, Advances in Neural Information Processing Systems 32: Annual Conference on Neural ...

  120. [128]

    Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, Raven Yuan, Tuney Zheng, Wei Pang, Xinrun Du, Yiming Liang, Yinghao Ma, Yizhi Li, Ziyang Ma, Bill Y. Lin, Emmanouil Benetos, Huan Yang, Junting Zhou, Ka...

  121. [129]

    Automathtext: Autonomous data selection with language models for mathematical texts

    Yifan Zhang, Yifan Luo, Yang Yuan, and Andrew Chi - Chih Yao. Automathtext: Autonomous data selection with language models for mathematical texts. CoRR, abs/2402.07625, 2024 b . doi:10.48550/ARXIV.2402.07625. URL https://doi.org/10.48550/arXiv.2402.07625

  122. [130]

    A survey of large language models

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian - Yun Nie, and Ji - Ro...

  123. [131]

    Opencodeinterpreter: Integrating code generation with execution and refinement

    Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, and Xiang Yue. Opencodeinterpreter: Integrating code generation with execution and refinement. In Lun - Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for C...

  124. [132]

    Programming every example: Lifting pre-training data quality like experts at scale

    Fan Zhou, Zengzhi Wang, Qian Liu, Junlong Li, and Pengfei Liu. Programming every example: Lifting pre-training data quality like experts at scale. CoRR, abs/2409.17115, 2024 a . doi:10.48550/ARXIV.2409.17115. URL https://doi.org/10.48550/arXiv.2409.17115

  125. [133]

    Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models

    Kun Zhou, Beichen Zhang, Jiapeng Wang, Zhipeng Chen, Wayne Xin Zhao, Jing Sha, Zhichao Sheng, Shijin Wang, and Ji - Rong Wen. Jiuzhang3.0: Efficiently improving mathematical reasoning by training small data synthesis models. CoRR, abs/2405.14365, 2024 b . doi:10.48550/ARXIV.24...

  126. [134]

    Yulan: An open-source large language model

    Yutao Zhu, Kun Zhou, Kelong Mao, Wentong Chen, Yiding Sun, Zhipeng Chen, Qian Cao, Yihan Wu, Yushuo Chen, Feng Wang, Lei Zhang, Junyi Li, Xiaolei Wang, Lei Wang, Beichen Zhang, Zican Dong, Xiaoxue Cheng, Yuhan Chen, Xinyu Tang, Yupeng Hou, Qiangqiang Ren, Xincheng Pang, Shufan...

  127. [135]

    Designing effective sparse expert models

    Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Designing effective sparse expert models. CoRR, abs/2202.08906, 2022. URL https://arxiv.org/abs/2202.08906

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.