REVIEW 3 major objections 4 minor 1 cited by
Domain-aware scaling laws recover data synergy and correctly rank optimal versus anti-optimal pretraining mixtures.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 07:22 UTC pith:H2HCYYFU
load-bearing objection Clean parametric extension of scaling laws that recovers usable domain synergies and correctly ranks new mixtures at small scale; the main open question is how far the low-dimensional observational axes generalize. the 3 major comments →
Domain-Aware Scaling Laws Uncover Data Synergy
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Observationally estimated first-order synergy coefficients correctly predict the performance ranking between predicted-optimal and predicted-anti-optimal pretraining mixtures on HumanEval, GSM8K and IFEval at both 30 M and 150 M scales, showing that domain-aware scaling laws recover actionable composition effects that classical domain-agnostic laws miss.
What carries the argument
Domain-aware scaling laws: classical Chinchilla-style power laws whose data term is rewritten so that each pretraining domain multiplies the effective data exponent by a benchmark-specific coefficient γ and, in the second-order extension, pairwise softmin co-occurrence terms σ produce benchmark-specific “bonus tokens” that vanish if either domain is absent.
Load-bearing premise
The low-dimensional mixture axes present in the public models are aligned enough with the true causal effects of domain co-occurrence that the fitted coefficients transfer to new controlled trainings.
What would settle it
Train models at the same scales on mixtures that the fitted coefficients rank as optimal and anti-optimal for a held-out benchmark; if the anti-optimal mixture systematically outperforms the optimal one, the claim that the observational coefficients are actionable fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes data synergy in LLM pretraining via domain-aware scaling laws. Starting from the Chinchilla form, it adds first-order domain o benchmark coefficients γ_j,k that modify the data scaling exponent per domain and benchmark (Eqs. 4–5), plus a shared second-order pairwise matrix Σ = {σ_kk′} that contributes softmin-gated “bonus” effective tokens when domains co-occur (Eq. 6 and Appendix A). Parameters are estimated observationally from 52 open-weight models spanning six families and eight domains, using Huber loss, ℓ1/ℓ2 regularization, a β prior, 5-fold CV, and bootstrap CIs. The FO and SO models raise held-out R² substantially over domain-agnostic baselines (median ~0.91 vs. much lower on code tasks) and recover sparse, sign-stable synergies (code–math and code–science positive; books interference). Actionability is tested by training new 30 M and 150 M models on FO-predicted optimal versus anti-optimal mixtures for HumanEval, GSM8K and IFEval; the ranking predicted by the observational γ holds at both scales (Table 1).
Significance. If the recovered coefficients transfer, the framework supplies a practical, low-cost tool for mixture design and data acquisition that goes beyond total-token scaling laws and additive mixture regressions. The observational estimation strategy, the explicit separation of first- and second-order effects, the bootstrap stability analysis, and especially the out-of-sample controlled trainings that confirm directional rankings are genuine strengths. The work therefore advances both the science of data composition and the engineering of pretraining corpora, even while remaining limited by the low-dimensional mixture geometry of public checkpoints and by the small scale of the validation runs.
major comments (3)
- [§4.3 / Table 1] §4.3 and Table 1: the decisive ranking validation is performed only at 30 M / 3 B tokens and 150 M / 5 B tokens, two orders of magnitude below the largest observational models (20 B). The paper assumes γ_j,k are scale-invariant, yet Appendix D only checks sign stability on the ≥1 B half of the observational set. Without either a larger-scale confirmation or a quantitative argument that the ranking gap persists under Chinchilla-optimal scaling, the claim that the observationally derived synergies are “actionable” for realistic pretraining remains under-supported.
- [Appendix D / §3] Appendix D and §3: two PCA components explain 95 % of the variance in the observed mixture vectors; the γ estimates are therefore composition sensitivities along those two axes rather than fully identified causal domain rates. The validation mixtures, while extreme (50 % math/code), still lie on the same low-dimensional span and do not probe intermediate weights, multi-domain interactions, or points outside the observed convex hull. The paper correctly labels the coefficients “associations,” yet the abstract and conclusion present them as transferable synergy estimates; a clearer delimitation of the domain of validity is required.
- [§4.2–4.3] §4.2–4.3: second-order coefficients σ_kk′ are estimated jointly and improve held-out fit on several benchmarks, yet the mixture-optimization experiments use only the first-order law. Consequently the claim that the framework “uncovers” second-order domain–domain synergy is only partially validated; it remains open whether including Σ would reorder the optimal/anti-optimal mixtures or further enlarge the observed gaps.
minor comments (4)
- [Figure 2] Figure 2 caption and surrounding text: the claim R² = 1.00 for the per-mixture exponent is almost certainly an over-fit on a tiny number of points; report the number of models per mixture and a leave-one-out or bootstrap R².
- [§2.3] Eq. (3) and the subsequent LSE forms: the entropy term −β_j H(u_i) is carried through but never ablated; a short note on its numerical contribution would help readers judge whether it is essential.
- [Appendix F] Table 8 (Appendix F): the “as published” rows for RegMix and Data Mixing Laws collapse because those methods were designed for fixed (N,D); the comparison is fair only after the “ext.” rows, which should be highlighted more prominently in the main text.
- [§2.4] Notation: zi,k is overloaded (first-order feature versus the softmin argument); a distinct symbol for the softmin inputs would improve readability.
Circularity Check
No load-bearing circularity: observational fits of γ/σ are tested by independent controlled trainings whose outcomes were never used in estimation.
full rationale
The paper's derivation chain is: (i) fit domain-aware scaling laws (Eqs. 4–6) to observational losses of 52 public checkpoints, recovering γ j,k and σ kk′; (ii) use the fitted first-order coefficients to score candidate mixtures via the derived objective Fj(c)=∑γ j,k h(ck) (Appendix G.1); (iii) train fresh 30 M/150 M models on the resulting predicted-optimal and anti-optimal mixtures; (iv) measure held-out BPB and confirm the ranking matches the prediction (Table 1, §4.3). Step (iv) is an out-of-sample empirical check whose targets were never seen by the fitter, so the ranking success is not forced by construction. Held-out R2 comparisons (Figs. 8–10) and bootstrap CIs are likewise standard predictive validation, not tautologies. The paper itself labels the coefficients associations rather than causal effects (§3) and notes the low-dimensional mixture space (Appendix D); those are validity caveats, not circular reductions of the form 'Eq. X = Eq. Y by definition' or 'fitted parameter renamed prediction.' No self-definitional loop, no uniqueness theorem imported from overlapping authors, and no ansatz smuggled via self-citation appear in the load-bearing path. Minor residual score of 1 reflects only that the observational design space is acknowledged to be limited, not that any claimed prediction collapses to its inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- per-benchmark data exponents β_j and domain modifiers γ_j,k
- shared pairwise synergy matrix Σ = {σ_kk′}
- Chinchilla scale coefficients (L∞, A, B, α) per benchmark
- softmin temperature τ and regularization strengths λ1, λ2, λβ
axioms (4)
- domain assumption Loss admits a log-sum-exp Chinchilla form whose data term can be multiplicatively rescaled by domain shares and pairwise softmin interactions.
- domain assumption The eight-domain taxonomy (books, code, encyclopedia, legal, math, Q&A, science, web) is a sufficient partition of pretraining sources.
- ad hoc to paper Rank-Gaussian transform of heterogeneous metrics (BPB, accuracy, perplexity) yields a common pseudo-loss scale suitable for joint fitting.
- ad hoc to paper Observational associations estimated under the low-dimensional mixture geometry of public models transfer directionally to new controlled trainings.
invented entities (2)
-
first-order domain→benchmark synergy coefficients γ_j,k
independent evidence
-
second-order domain-domain synergy matrix σ_kk′ and associated “bonus” effective tokens
no independent evidence
read the original abstract
Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings repeatedly show that combining datasets from different domains yields nontrivial interactions. For instance, adding code improves mathematical reasoning, while certain mixtures introduce interference that reduces model performance. We refer to these effects collectively as data synergy, where the contribution of multiple domains exceeds or falls short of the sum of their isolated contributions. In this work, we formalize and quantify data synergy in language model pretraining. Leveraging observational variation across open-weight LLMs with diverse pretraining mixtures, we estimate both direct domain-to-benchmark synergy (how one domain contributes to performance on another) and a second-order domain-domain synergy (capabilities that require co-occurrence of multiple domains). Our framework improves predictive accuracy over domain-agnostic scaling laws and recovers stable synergy estimates. We validate these estimates by training models on predicted optimal and predicted anti-optimal mixtures and confirm that our synergy estimates correctly predict performance rankings.
Figures
Forward citations
Cited by 1 Pith paper
-
Bridging Compute- and Data-Optimal Pretraining
Pretraining loss obeys a single law in which repeated or paraphrased tokens count as η(N, data-per-parameter, expansion-ratio) fresh tokens, with total effective data saturating as derived tokens grow.
Reference graph
Works this paper leans on
-
[1]
To code, or not to code? exploring impact of code in pre-training.arXiv preprint arXiv:2408.10914,
Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. To code, or not to code? exploring impact of code in pre-training.arXiv preprint arXiv:2408.10914,
-
[2]
Program synthesis with large language models.arXiv preprint arXiv:2108.07732,
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models.arXiv preprint arXiv:2108.07732,
-
[3]
Llemma: An open language model for mathematics.arXiv preprint arXiv:2310.10631,
Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics.arXiv preprint arXiv:2310.10631,
-
[4]
If you use this software, please cite it using these metadata
URL https://doi.org/10.5281/zenodo.5297715. If you use this software, please cite it using these metadata. Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. Gpt-neox-20b: An open- source autoregressive language model. InProceedings of BigScience Episode# 5–...
-
[5]
Evalu- ating large language models trained on code.arXiv preprint arXiv:2107.03374,
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evalu- ating large language models trained on code.arXiv preprint arXiv:2107.03374,
-
[6]
Mayee F Chen, Michael Y Hu, Nicholas Lourie, Kyunghyun Cho, and Christopher Ré. Aioli: A unified optimization framework for language model data mixing.arXiv preprint arXiv:2411.05735, 2024a. Yangyi Chen, Binxuan Huang, Yifan Gao, Zhengyang Wang, Jingfeng Yang, and Heng Ji. Scaling laws for predicting downstream performance in llms.arXiv preprint arXiv:241...
-
[7]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge.arXiv preprint arXiv:1803.05457,
-
[8]
Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,
-
[9]
URL https://zenodo.org/ records/12608602. Albert Ge, Tzu-Heng Huang, John Cooper, Avi Trost, Ziyi Chu, Satya Sai Srinath Namburi GNVV , Ziyang Cai, Kendall Park, Nicholas Roberts, and Frederic Sala. R&b: Domain regrouping and data mixture balancing for efficient foundation model training.arXiv preprint arXiv:2505.00358,
-
[10]
URL https://github.com/openlm-research/open_llama. Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Acceler- ating the science of language models.arXiv preprint arXiv:2402.00838,
-
[11]
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. Studying large language model generalization with influence functions.arXiv preprint arXiv:2308.03296,
-
[12]
Jiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao, and Fei Tan. Cmr scaling law: Predicting critical mixture ratios for continual pre-training of language models.arXiv preprint arXiv:2407.17467,
-
[13]
Training compute-optimal large language models.arXiv preprint arXiv:2203.15556,
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models.arXiv preprint arXiv:2203.15556,
-
[14]
Datamodels: Predicting predictions from training data.arXiv preprint arXiv:2202.00622,
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. Datamodels: Predicting predictions from training data.arXiv preprint arXiv:2202.00622,
-
[15]
Jikai Jin, Vasilis Syrgkanis, Sham Kakade, and Hanlin Zhang. Discovering hierarchical latent capabilities of language models via causal representation learning.arXiv preprint arXiv:2506.10378,
-
[16]
Autoscale: Scale-aware data mixing for pre-training llms.arXiv preprint arXiv:2407.20177,
Feiyang Kang, Yifan Sun, Bingbing Wen, Si Chen, Dawn Song, Rafid Mahmood, and Ruoxi Jia. Autoscale: Scale-aware data mixing for pre-training llms.arXiv preprint arXiv:2407.20177,
-
[17]
Scaling laws for neural language models.arXiv preprint arXiv:2001.08361,
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv:2001.08361,
Pith/arXiv arXiv 2001
-
[18]
Najoung Kim, Sebastian Schuster, and Shubham Toshniwal. Code pretraining improves entity tracking abilities of language models.arXiv preprint arXiv:2405.21068,
-
[19]
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations. InProceedings of the 2017 conference on empirical methods in natural language processing, pp. 785–794,
2017
-
[20]
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Guha, Sedrick Scott Keh, Kushal Arora, et al. Datacomp-lm: In search of the next generation of training sets for language models.Advances in Neural Information 11 Processing Systems, 37:14200–14282, 2024a. Mingxin Li, Zhijie Nie, Yanzhao Zhang, Dingk...
-
[21]
Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, et al. The data provenance initiative: A large scale audit of dataset licensing & attribution in ai.arXiv preprint arXiv:2310.16787,
-
[22]
Shayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I Hsu, Isaac Caswell, Alex Pent- land, Sercan Arik, Chen-Yu Lee, Sayna Ebrahimi, et al. Atlas: Adaptive transfer scaling laws for multilingual pretraining, finetuning, and decoding the curse of multilinguality. arXiv preprint arXiv:2510.22037,
-
[23]
Zimu Lu, Aojun Zhou, Ke Wang, Houxing Ren, Weikang Shi, Junting Pan, Mingjie Zhan, and Hongsheng Li. Mathcoder2: Better math reasoning from continued pretraining on model-translated mathematical code.arXiv preprint arXiv:2410.08196,
-
[24]
At which training stage does code data help llms reasoning?arXiv preprint arXiv:2309.16298,
Yingwei Ma, Yue Liu, Yue Yu, Yuanliang Zhang, Yu Jiang, Changjian Wang, and Shan- shan Li. At which training stage does code data help llms reasoning?arXiv preprint arXiv:2309.16298,
-
[25]
Ian Magnusson, Nguyen Tai, Ben Bogin, David Heineman, Jena D Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu, Dirk Groeneveld, Oyvind Tafjord, et al. Datadecide: How to predict best pretraining data with small experiments.arXiv preprint arXiv:2504.11393,
-
[26]
Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843,
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843,
-
[27]
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. The lambada dataset: Word prediction requiring a broad discourse context.arXiv preprint arXiv:1606.06031,
-
[28]
Rylan Schaeffer, Noam Levi, Brando Miranda, and Sanmi Koyejo. Pretraining scaling laws for generative evaluations of language models.arXiv preprint arXiv:2509.24012,
-
[29]
Scaling laws for optimal data mixtures.arXiv preprint arXiv:2507.09404,
Mustafa Shukor, Louis Bethune, Dan Busbridge, David Grangier, Enrico Fini, Alaaeldin El-Nouby, and Pierre Ablin. Scaling laws for optimal data mixtures.arXiv preprint arXiv:2507.09404,
-
[30]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. Dolma: An open corpus of three trillion tokens for language model pretraining research.arXiv preprint 12 arXiv:2402.00159,
-
[31]
Tristan Thrush, Christopher Potts, and Tatsunori Hashimoto
URL https://arxiv.org/abs/2501.00656. Tristan Thrush, Christopher Potts, and Tatsunori Hashimoto. Improving pretraining data using perplexity correlations.arXiv preprint arXiv:2409.05816,
-
[32]
Jiapeng Wang, Changxin Tian, Kunlong Chen, Ziqi Liu, Jiaxin Mao, Wayne Xin Zhao, Zhiqiang Zhang, and Jun Zhou. Mergemix: Optimizing mid-training data mixtures via learnable model merging.arXiv preprint arXiv:2601.17858,
-
[33]
Data mixing laws: Optimizing data mixtures by predicting language modeling performance
Jiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan, Yunhua Zhou, and Xipeng Qiu. Data mixing laws: Optimizing data mixtures by predicting language modeling performance. arXiv preprint arXiv:2403.16952,
-
[34]
Huimu Yu, Xing Wu, Haotian Xu, Debing Zhang, and Songlin Hu. Codepmp: Scal- able preference model pretraining for large language model reasoning.arXiv preprint arXiv:2410.02229,
-
[35]
Group- level data selection for efficient pretraining.arXiv preprint arXiv:2502.14709,
Zichun Yu, Fei Peng, Jie Lei, Arnold Overwijk, Wen-tau Yih, and Chenyan Xiong. Group- level data selection for efficient pretraining.arXiv preprint arXiv:2502.14709,
-
[36]
Hellaswag: Can a machine really finish your sentence?arXiv preprint arXiv:1905.07830,
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence?arXiv preprint arXiv:1905.07830,
Pith/arXiv arXiv 1905
-
[37]
Junhao Zheng, Qianli Ma, Zhen Liu, Binquan Wu, and Huawen Feng. Beyond anti- forgetting: Multimodal continual instruction tuning with positive forward transfer.arXiv preprint arXiv:2401.09181,
-
[38]
Instruction-following evaluation for large language models.arXiv preprint arXiv:2311.07911,
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-following evaluation for large language models.arXiv preprint arXiv:2311.07911,
-
[39]
Here we give the per-group scales and checkpoint counts
summarizes the model set. Here we give the per-group scales and checkpoint counts. Our model set spans six open-weight model groups with different scales and training mixtures: GPT-Neo/J/NeoX (Black et al., 2021; Wang & Komatsuzaki, 2021; Black et al.,
2021
-
[40]
Six domains (Books, Code, Encyclopedia, Legal, Science, Web) match DPI source domains
and the Level 1 categories of Essential-Web v1.0 (Essential AI, 2025). Six domains (Books, Code, Encyclopedia, Legal, Science, Web) match DPI source domains. We addMathfor the explicit math datasets (DM Mathematics, OpenWebMath, Algebraic Stack) that match Essential-Web’s math category, andQ&Afor sources that are exclusively questions and answers in the m...
2025
-
[41]
and MBPP (Austin et al., 2021), mathematical reasoning with GSM8K (Cobbe et al., 2021), science reasoning with ARC-Challenge (Clark et al., 2018), commonsense inference with HellaSwag (Zellers et al., 2019), reading comprehension with RACE (Lai et al., 2017), open-domain QA with TriviaQA (Joshi et al., 2017), cloze completion withLAMBADA (Paperno et al., ...
2021
-
[42]
All evaluations use the lm-evaluation-harness (Gao et al., 2024)
and C4 (Raffel et al., 2019). All evaluations use the lm-evaluation-harness (Gao et al., 2024). Because our model set includes many small models (70M–1B), metrics such as pass@1 and accuracy are close to zero for some generative tasks. For example, 55% of the models score exactly zero on HumanEval pass@1, and 52% on MBPP. With the majority of observations...
2019
-
[43]
Second, our second-order model introduces an explicit shared pairwise interaction matrix Σ={σ kk′ } that captures co-occurrence effects that vanish when either domain is absent, while none of the previous functional forms include such a term. Third, all of the above are fit from controlledsmall-scale training runs deliberately swept over mixture and scale...
2022
-
[44]
Appendix G.1 derives the objective used to choose the validation mixtures, and Appendix G.2 provides more details for this experiment
0.024 0.014−0.178 RegMix (as published) (Liu et al., 2024)−0.080−0.074−0.185 G Mixture Optimization and Validation This appendix expands on the mixture validation experiments of Section 4.3. Appendix G.1 derives the objective used to choose the validation mixtures, and Appendix G.2 provides more details for this experiment. G.1 Mixture Optimization Deriva...
2024
-
[45]
The 30M model uses dmodel = 256, 8 heads, and 8 layers, while the 150M model uses dmodel = 640, 10 heads, and 12 layers
framework on a single A100 GPU per run, at two scales: 30M-parameter models on 3B tokens and 150M-parameter models on 5B tokens. The 30M model uses dmodel = 256, 8 heads, and 8 layers, while the 150M model uses dmodel = 640, 10 heads, and 12 layers. Both use sequence length 2048 and AdamW with Chinchilla scaled learning rates. Training data comes from the...
2048
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.