Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Compute Requirements for Algorithmic Innovation in Frontier AI Models

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Even strict compute caps would still leave room for half of AI's algorithmic innovations, a catalog of 36 techniques suggests.

desk verdict First useful dataset on compute costs for algorithmic innovations, but the 'half of innovations' count does not establish the policy conclusion about progress. read the letter →

arxiv 2507.10618 v1 pith:4I2VG32M submitted 2025-07-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords algorithmicprogresscomputegovernancecapspretraininginnovationslanguagemodelsscalinglawsFLOPestimationhardwarecapacity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish whether restrictions on computing resources can slow the development of new AI algorithms. It catalogs 36 pre-training algorithmic innovations implemented in Llama 3 and DeepSeek-V3, estimates the FLOP and hardware capacity used to develop each, and then asks what fraction would remain developable under various compute caps. It finds that even very stringent caps, such as limiting total operations to GPT-2's training compute or hardware to 8 H100 GPUs, would still have allowed roughly half of the cataloged innovations. If this is right, compute-focused governance would need to be paired with other levers to manage algorithmic progress.

What carries the argument

The load-bearing object is a hand-built catalog of 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3, each with an estimated total FLOP and hardware capacity in TFLOP/s drawn from the experiments in its original paper. The counterfactual analysis plots the cumulative fraction of innovations whose estimated requirements fall below a given FLOP cap or hardware cap, using GPT-2 training compute and H100 counts as reference points. The distinction between total operations and hardware capacity matters because innovations like parallelization techniques may use negligible FLOP but require large clusters, so the two cap types are analyzed separately. Roughly 25 percent of the cataloged innovations used negligible FLOP.

What would settle it

A credible test would take a sample of the 36 innovations and compare the reported paper compute against audited internal development records, including failed runs; if the true development compute is consistently several times higher, the caps' allowed fraction would drop below half.

Watch

Extended reading notes

Core claim

The central claim is that compute caps alone are unlikely to dramatically slow AI algorithmic progress. The paper builds a dataset of 36 pre-training innovations from two open frontier model families, assigns each an estimated development FLOP and TFLOP/s based on the experiments in its original paper, and shows that non-negligible innovations grow at roughly 2.5 times per year in FLOP and 2.1 times per year in hardware capacity. Under counterfactual caps at GPT-2-level FLOP or 8 H100s, about half of the cataloged innovations still fall below the cap. The author reads this as evidence that algorithmic progress has a low near-term compute threshold, so restricting compute alone is unlikely to stop or dramatically slow the development of new pretraining algorithms.

Load-bearing premise

The analysis assumes the compute reported in each original paper is close to the compute actually needed to develop that innovation, even though failed experiments and validation at scale are omitted from the papers.

Editorial extensions

If this is right

  • If the estimates hold, regulators cannot rely on compute caps as a standalone brake on algorithmic progress; half of recent innovations would survive even very tight caps.
  • Export controls that limit rival states to modest hardware would probably not stop algorithmic improvement, since existing hardware is already sufficient for many innovations.
  • The doubling trend implies that caps set at fixed 2025-style levels may become increasingly binding over time, but until then most innovations remain below them.
  • Innovations that save compute at training time remain discoverable under caps precisely because their development does not require large training runs.
  • Compute governance appears more credible when paired with legal or institutional measures, as the paper itself concludes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's estimates are lower bounds on real R&D compute because failed experiments and scale-up validation are excluded; a systematic calibration of reported versus true compute could shift the cap curve downward.
  • Under enforced scarcity, researchers would likely shift toward low-FLOP innovations, so the realized innovation rate under caps might be higher than the static catalog suggests.
  • The catalog covers only open pre-training innovations; post-training and closed-lab innovations could have different compute profiles, so the policy conclusion may not transfer to those domains.
  • A natural testable extension is to apply the same catalog method to post-training and reasoning innovations, where compute requirements are currently lower but rising.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper catalogs 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3, estimates for each the total FLOP and hardware capacity (TFLOP/s) used in the introducing paper, and analyzes how these estimates evolve over time. The authors report that non-negligible-FLOP innovations have grown at roughly 2.53x per year in total operations and 2.14x per year in hardware capacity. They then use the catalog to estimate what fraction of innovations would remain available under compute caps, concluding that even stringent caps (GPT-2-scale FLOP or 8 H100s) would still allow about half of the cataloged innovations, and hence that compute caps alone are unlikely to dramatically slow AI algorithmic progress. The paper is transparent about limitations, including omitted failed experiments, exclusion of proprietary labs, and the absence of impact weighting.

Significance. If the central conclusion were established, this would be a valuable contribution to compute-governance debates: it provides the first systematic catalog of compute requirements for algorithmic innovations in open frontier models, along with concrete trend estimates and a falsifiable 2028 projection. The Appendix A dataset is transparent and will be useful for future work on algorithmic progress. The paper is also commendably candid in Section 5.1 about the many ways its estimates could be biased. However, the headline policy inference moves from an unweighted count of innovations to a claim about the rate of algorithmic progress, and the manuscript's own Section 5.2 defers impact weighting to future work. That gap is load-bearing, so the central claim is not yet supported as stated.

major comments (4)
  1. [§4, Abstract] The headline conclusion equates the number of cataloged innovations below a cap with the rate of algorithmic progress. Figure 3 reports cumulative counts, but no impact weighting is provided; Section 5.2 explicitly lists 'Quantifying Innovation Impact (CEG)' as future work. The blocked set can plausibly contain the highest-impact innovations: under a GPT-2-scale FLOP cap, the blocked set includes Chinchilla scaling laws, DeepSeekMoE, MLA, and FP8-LM, and under an 8-H100 hardware cap it includes ZeRO, tensor parallelism, and FSDP. If these carry most of the compute-equivalent gain, progress could slow dramatically even though half the count is available. The count statistic is therefore insufficient for the progress-based conclusion; the text should either provide an impact-weighted analysis or restrict the claim to 'half of the cataloged innovations.'
  2. [§5.1, §4] The counterfactual analysis treats the compute reported in the introducing paper as the compute required to develop the innovation. Section 5.1 acknowledges that reported costs omit failed experiments and preliminary explorations, and that validation at scale may require much more compute, making the estimates lower bounds on actual R&D compute. The Section 4 caveat that researchers could be more efficient under a cap is a different bias and does not cancel this one. If actual development compute is systematically higher than reported, caps would block more than half of the innovations. The paper should bound or quantify this bias, for example through the researcher interviews suggested in Section 5.2 or through sensitivity analysis that inflates reported compute by plausible factors.
  3. [Table 2, Figure 3] The cap fractions are deterministic functions of point estimates in Table 2, yet those estimates are reported to several significant figures with no uncertainty and come from heterogeneous sources, including personal correspondence. Several entries have missing TFLOP/s values (marked '—'), and it is unclear how Figure 3 (bottom) treats these missing values when computing the fraction below a hardware cap. A small number of mis-estimated entries could move the 'half' result. The paper should clarify the handling of missing values and add sensitivity analysis around the point estimates, for example by showing how the cap fractions change under factor-of-2 or factor-of-10 perturbations.
  4. [§1, §5.1] The sample is restricted to 36 innovations in two open model families, and Section 5.1 acknowledges the exclusion of proprietary labs. Yet the Abstract and Section 4 draw conclusions about 'AI algorithmic progress' in general. Since the most compute-intensive innovations may occur in closed labs, the representativeness of the open-model catalog is load-bearing for the general policy claim. The manuscript should either justify that the open-model sample is representative of frontier algorithmic progress or explicitly restrict the conclusion to the population of open-model pre-training innovations.
minor comments (5)
  1. [§2] The phrase 'terraFLOP/s' should be 'teraFLOP/s'.
  2. [§4] The sentence 'These caps are would still be fairly low' contains a grammatical error and should read 'These caps would still be fairly low.'
  3. [Figure 1 caption] The caption should state explicitly that the trend line is fit only to innovations that do not use negligible FLOP, and should clarify what the shaded area represents (the text says 95% CI but does not say whether it is a confidence band for the mean or prediction interval).
  4. [Table 2] The 'Math equiv' column uses a star symbol that is never defined in the table caption; please define it and explain how 'Negligible' FLOP and missing TFLOP/s entries are treated in Figure 3.
  5. [§4] The sentence beginning 'A further implication is that US-led export controls...' goes beyond the evidence presented, since export controls target specific actors and involve different enforcement and verification dynamics than the domestic caps modeled here; consider softening this claim or adding supporting reasoning.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the cap analysis is a transparent cumulative count of the paper's own catalog, not a hidden fit or self-citation chain.

full rationale

The paper's derivation chain is self-contained and non-circular. The central empirical result — that half of the cataloged innovations fall below the GPT-2 FLOP cap and the 8-H100 hardware cap — is a direct cumulative distribution of the paper's own dataset, explicitly presented as a preliminary measure in Section 4 ('As a preliminary measure of the impact of compute caps on algorithmic progress, we calculate what fraction of the cataloged innovations fall below a given FLOP or hardware cap'). No parameter is fitted to a target and then relabeled as a prediction; the trend regressions in Section 3 are descriptive summaries and are not used to construct the cap conclusion. The paper's own limitations section candidly notes that the reported compute is a lower bound on actual R&D compute and that impact weighting (CEG) is left to future work, but those are evidentiary and interpretive gaps, not circularity. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation. The inference from 'half of the cataloged innovations available' to 'compute caps alone are unlikely to dramatically slow AI algorithmic progress' is an extrapolation that may be debated on representativeness or impact-weighting grounds, but it is not equivalent to its inputs by construction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on four domain assumptions inherited from the measurement design: reported published compute as a proxy for required compute, representativeness of the open-model sample, independence of innovations in the cap counterfactual, and validity of log-linear extrapolation for the 2028 projection. There are no invented physical or mathematical entities.

free parameters (2)
  • FLOP growth rate (log-linear trend slope) = 2.53x per year (95% CI 1.86-3.38)
    Fitted by regression to the non-negligible-FLOP catalog entries in Figure 1; supports the 'doubling each year' trend claim, not the cap counterfactual.
  • Hardware capacity growth rate = 2.14x per year (95% CI 1.44-2.76)
    Fitted by regression to hardware capacity data in Figure 1; supports the claim that innovation compute demand is rising faster than compute price-performance.
assumptions (4)
  • domain assumption The compute reported in the paper that introduced an innovation measures the compute required to develop it.
    Used throughout Section 2 and the cap analysis in Section 4; Section 5.1 acknowledges reported compute is a lower bound on actual compute because failed experiments are omitted.
  • domain assumption The 36 innovations used in Llama 3 and DeepSeek-V3 are a representative sample of pre-training algorithmic innovation.
    Section 2 chooses these two open model families for transparency; Section 5.1 notes proprietary models and post-training innovations are excluded, so generalizing the fraction below caps to all algorithmic progress is an assumption.
  • domain assumption An innovation is 'available' under a cap if its estimated development compute is below the cap, independent of other innovations and prior lineage.
    Section 4 builds the cumulative count this way; Section 5.1, 'Ignoring Research Lineage', notes cumulative upstream compute is not counted.
  • domain assumption Log-linear extrapolation of the observed trend to 2028 is valid for forecasting future median compute requirements.
    Used in Section 4 to project median hardware capacity of 10^6 TFLOP/s and median total operations of 10^24 FLOP by 2028; this projection is not needed for the main cap counterfactual.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Compute Requirements for Algorithmic Innovation in Frontier AI Models." pith.science (2026). https://pith.science/paper/4I2VG32M

@misc{pith2026250710618,
  author       = {Pith},
  title        = {Pith review of: Compute Requirements for Algorithmic Innovation in Frontier AI Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4I2VG32M}},
  note         = {Machine review of arXiv:2507.10618}
}
read the original abstract

Algorithmic innovation in the pretraining of large language models has driven a massive reduction in the total compute required to reach a given level of capability. In this paper we empirically investigate the compute requirements for developing algorithmic innovations. We catalog 36 pre-training algorithmic innovations used in Llama 3 and DeepSeek-V3. For each innovation we estimate both the total FLOP used in development and the FLOP/s of the hardware utilized. Innovations using significant resources double in their requirements each year. We then use this dataset to investigate the effect of compute caps on innovation. Our analysis suggests that compute caps alone are unlikely to dramatically slow AI algorithmic progress. Even stringent compute caps -- such as capping total operations to the compute used to train GPT-2 or capping hardware capacity to 8 H100 GPUs -- could still have allowed for half of the cataloged innovations.

Figures

Figures reproduced from arXiv: 2507.10618 by the authors.

Figure 1
Figure 1. Top: The total operations (in FLOP) for the different algorithmic innovations over time, colored by their category. The triangles indicate innovations which used negligible FLOP in their experiments. The black line shows the trend, and the shaded area indicates the 95% CI for this trend. This trend line is only for innovations that did not use negligible compute. Bottom: The hardware capacity (in TFLOP/s) for the di… view at source ↗
Figure 2
Figure 2. Total operations used for experiments versus hardware capacity for the different algorithmic innovations. Compute used is correlated with hardware capacity. This graph only includes innovations which used non-negligible FLOP. 4. Impact of Compute Caps Compute caps may reduce the rate of algorithmic progress. Along with the direct impact of preventing the develop￾ment of dangerous AI systems, caps may also increase t… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How to Catch a GPU: A Taxonomy of Verification and Enforcement Mechanisms for International AI Agreements

    cs.CY 2026-06 conditional novelty 6.0 of 10

    Verification of international AI agreements will fail first at detecting hidden compute facilities, around the 10,000-H100-equivalent scale, before other enforcement mechanisms break.

Reference graph

Works this paper leans on

61 extracted references · 25 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead

    Aguirre, A. Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead . arXiv preprint arXiv:2311.09452, 2025. URL https://arxiv.org/abs/2311.09452v4

  3. [3]

    d., Zemlyanskiy, Y., Lebron, F., and Sanghai, S

    Ainslie, J., Lee-Thorp, J., Jong, M. d., Zemlyanskiy, Y., Lebron, F., and Sanghai, S. GQA : Training Generalized Multi - Query Transformer Models from Multi - Head Checkpoints . December 2023. URL https://openreview.net/forum?id=hmOwOZWzYE

  4. [4]

    Singe: leveraging warp specialization for high performance on GPUs

    Bauer, M., Treichler, S., and Aiken, A. Singe: leveraging warp specialization for high performance on GPUs . SIGPLAN Not., 49 0 (8): 0 119--130, February 2014. ISSN 0362-1340. doi:10.1145/2692916.2555258. URL https://doi.org/10.1145/2692916.2555258

  5. [5]

    Efficient Training of Language Models to Fill in the Middle , July 2022

    Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. Efficient Training of Language Models to Fill in the Middle , July 2022. URL http://arxiv.org/abs/2207.14255. arXiv:2207.14255 [cs]

  6. [6]

    and Aarne, O

    Brass, A. and Aarne, O. Location verification for ai chips. Issue brief, Institute for AI Policy and Strategy (IAPS), April 2024. URL https://www.iaps.ai/research/location-verification-for-ai-chips

  7. [7]

    Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models

    Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. arXiv preprint arXiv:2401.06066, 2024

  8. [8]

    Y., Ermon, S., Rudra, A., and Ré, C

    Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. FLASHATTENTION : fast and memory-efficient exact attention with IO -awareness. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS '22, pp.\ 16344--16359, Red Hook, NY, USA, November 2022. Curran Associates Inc. ISBN 978-1-71387-108-8

Show all 61 references
  1. [9]

    Ai capabilities can be significantly improved without expensive retraining

    Davidson, T., Denain, J.-S., Villalobos, P., and Bas, G. Ai capabilities can be significantly improved without expensive retraining. arXiv preprint arXiv:2312.07413, 2023

  2. [10]

    DeepSeek - Coder - V2 : Breaking the Barrier of Closed - Source Models in Code Intelligence , June 2024

    DeepSeek-AI, Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., Zeng, W., Bi, X., Gu, Z., Xu, H., Dai, D., Dong, K., Zhang, L., Piao, Y., Gou, Z., Xie, Z., Hao, Z., Wang, B., Song, J., Chen, D., Xie, X., Guan, K., You, Y., Liu, A., Du, Q.,...

  3. [11]

    LLM .int8(): 8-bit matrix multiplication for transformers at scale

    Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. LLM .int8(): 8-bit matrix multiplication for transformers at scale. In Proceedings of the 36th International Conference on Neural Information Processing Systems , NIPS '22, pp.\ 30318--30332, Red Hook, NY, USA, November...

  4. [12]

    A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., and Shen, J

    Douillard, A., Feng, Q., Rusu, A. A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., and Shen, J. Diloco: Distributed low-communication training of language models. arXiv preprint arXiv:2311.08105, 2023

  5. [13]

    Data on machine learning hardware”, 10 2024 a

    Epoch AI . Data on machine learning hardware”, 10 2024 a . URL https://epoch.ai/data/machine-learning-hardware. Accessed: 2025-04-28

  6. [14]

    Data on notable ai models, 6 2024 b

    Epoch AI . Data on notable ai models, 6 2024 b . URL https://epoch.ai/data/notable-ai-models. Accessed: 2025-05-11

  7. [15]

    Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G

    Gloeckle, F., Idrissi, B. Y., Rozière, B., Lopez-Paz, D., and Synnaeve, G. Better & faster large language models via multi-token prediction. In Proceedings of the 41st International Conference on Machine Learning , volume 235 of ICML '24 , pp.\ 15706--15734, Vienna, Austria, J...

  8. [16]

    The llama 3 herd of models, 2024

    Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. The llama 3 herd of models, 2024

  9. [17]

    H., Ivison, H., Magnusson, I., Wang, Y., et al

    Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al. Olmo: Accelerating the science of language models. arXiv preprint arXiv:2402.00838, 2024

  10. [18]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning

    Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025

  11. [19]

    and Koessler, L

    Heim, L. and Koessler, L. Training compute thresholds: Features and functions in ai regulation. arXiv preprint arXiv:2405.10799, 2024

  12. [20]

    A., and Zilberman, N

    Heim, L., Fist, T., Egan, J., Huang, S., Zekany, S., Trager, R., Osborne, M. A., and Zilberman, N. Governing through the cloud: The intermediary role of compute providers in ai regulation. arXiv preprint arXiv:2403.08501, 2024

  13. [21]

    C., Atkinson, D., Thompson, N., and Sevilla, J

    Ho, A., Besiroglu, T., Erdil, E., Owen, D., Rahman, R., Guo, Z. C., Atkinson, D., Thompson, N., and Sevilla, J. Algorithmic progress in language models, 2024

  14. [22]

    Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyal...

  15. [23]

    X., Chen, D., Lee, H., Ngiam, J., Le, Q

    Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z. GPipe : efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33rd International Conference on Neural Information Proc...

  16. [24]

    Openai o1 system card

    Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024

  17. [25]

    M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., et al

    Jaghouar, S., Ong, J. M., Basra, M., Obeid, F., Straube, J., Keiblinger, M., Bakouch, E., Atkins, L., Panahi, M., Goddard, C., et al. Intellect-1 technical report. arXiv preprint arXiv:2412.01152, 2024

  18. [26]

    A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

    Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B. Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems, 5: 0 341--353, 2023

  19. [27]

    and Richardson, J

    Kudo, T. and Richardson, J. SentencePiece : A simple and language independent subword tokenizer and detokenizer for Neural Text Processing . In Blanco, E. and Lu, W. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing : System Demonst...

  20. [28]

    Kulp, G., Gonzales, D., Smith, E., Heim, L., Puri, P., Vermeer, M. J. D., and Winkelman, Z. Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090. RAND Corporation...

  21. [29]

    Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., et al. T\"ulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024

  22. [30]

    Breadth- First Pipeline Parallelism

    Lamy-Poirier, J. Breadth- First Pipeline Parallelism . Proceedings of Machine Learning and Systems, 5: 0 48--67, March 2023. URL https://proceedings.mlsys.org/paper_files/paper/2023/hash/24e845415c1486dd2d582a9d639237f9-Abstract-mlsys2023.html

  23. [31]

    Y., Bansal, H., Guha, E., Keh, S

    Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S. Y., Bansal, H., Guha, E., Keh, S. S., Arora, K., et al. Datacomp-lm: In search of the next generation of training sets for language models. Advances in Neural Information Processing Systems, 37: 0 14200--14282, 2024

  24. [32]

    DeepSeek - V2 : A Strong , Economical , and Efficient Mixture -of- Experts Language Model , 2024 a

    Liu, A., Feng, B., Wang, B., Wang, B., Liu, B., Zhao, C., Dengr, C., Ruan, C., Dai, D., Guo, D., et al. DeepSeek - V2 : A Strong , Economical , and Efficient Mixture -of- Experts Language Model , 2024 a

  25. [33]

    DeepSeek-V3 Technical Report , 2024 b

    Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., Ruan, C., et al. DeepSeek-V3 Technical Report , 2024 b

  26. [34]

    RingAttention with Blockwise Transformers for Near - Infinite Context

    Liu, H., Zaharia, M., and Abbeel, P. RingAttention with Blockwise Transformers for Near - Infinite Context . October 2023. URL https://openreview.net/forum?id=WsRHpHH4s0

  27. [35]

    and Hutter, F

    Loshchilov, I. and Hutter, F. Decoupled Weight Decay Regularization . September 2018. URL https://openreview.net/forum?id=Bkg6RiCqY7

  28. [36]

    A Narrow Path , December 2024

    Miotti, A., Bilge, T., Kasten, D., and Newport, J. A Narrow Path , December 2024. URL https://www.narrowpath.co/

  29. [37]

    Efficient large-scale language model training on GPU clusters using megatron- LM

    Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M. Efficient large-scale language model training on GPU clusters using megatron- LM . In Proceedings ...

  30. [38]

    8-bit Numerical Formats for Deep Neural Networks , June 2022

    Noune, B., Jones, P., Justus, D., Masters, D., and Luschi, C. 8-bit Numerical Formats for Deep Neural Networks , June 2022. URL http://arxiv.org/abs/2206.02915. arXiv:2206.02915 [cs]

  31. [39]

    2 olmo 2 furious

    OLMo Team , Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., et al. 2 olmo 2 furious. arXiv preprint arXiv:2501.00656, 2024

  32. [40]

    Training language models to follow instructions with human feedback

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 27730--27744, 2022

  33. [41]

    FP8 - LM : Training FP8 Large Language Models , December 2023

    Peng, H., Wu, K., Wei, Y., Zhao, G., Yang, Y., Liu, Z., Xiong, Y., Yang, Z., Ni, B., Hu, J., Li, R., Zhang, M., Li, C., Ning, J., Wang, R., Zhang, Z., Liu, S., Chau, J., Hu, H., and Cheng, P. FP8 - LM : Training FP8 Large Language Models , December 2023. URL http://arxiv.org/a...

  34. [42]

    Interim report: Mechanisms for flexible hardware-enabled guarantees

    Petrie, J., Aarne, O., Ammann, N., and Dalrymple, D. Interim report: Mechanisms for flexible hardware-enabled guarantees. Technical report, 8 2024

  35. [43]

    Zero Bubble ( Almost ) Pipeline Parallelism

    Qi, P., Wan, X., Huang, G., and Lin, M. Zero Bubble ( Almost ) Pipeline Parallelism . October 2023. URL https://openreview.net/forum?id=tuzTN0eIO5

  36. [44]

    Rabe, M. N. and Staats, C. Self-attention does not need o (n^2) memory, 2021

  37. [45]

    ZeRO : memory optimizations toward training trillion parameter models

    Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. ZeRO : memory optimizations toward training trillion parameter models. In Proceedings of the International Conference for High Performance Computing , Networking , Storage and Analysis , SC '20, pp.\ 1--16, Atlanta, Georgia, ...

  38. [46]

    Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y

    Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y. ZeRO - Offload : Democratizing Billion - Scale Model Training . January 2021. URL https://openreview.net/forum?id=qXFQtGMHRa

  39. [47]

    K., Ngo, R., Pilz, K., et al

    Sastry, G., Heim, L., Belfield, H., Anderljung, M., Brundage, M., Hazell, J., O'Keefe, C., Hadfield, G. K., Ngo, R., Pilz, K., et al. Computing power and the governance of artificial intelligence. arXiv preprint arXiv:2402.08797, 2024

  40. [48]

    and Thiergart, L

    Scher, A. and Thiergart, L. Mechanisms to Verify International Agreements About AI Development , November 2024. URL https://techgov.intelligence.org/research/mechanisms-to-verify-international-agreements-about-ai-development

  41. [49]

    Neural Machine Translation of Rare Words with Subword Units

    Sennrich, R., Haddow, B., and Birch, A. Neural Machine Translation of Rare Words with Subword Units . In Erk, K. and Smith, N. A. (eds.), Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pp.\ 1715--1725, Berlin...

  42. [50]

    GLU Variants Improve Transformer , February 2020

    Shazeer, N. GLU Variants Improve Transformer , February 2020. URL http://arxiv.org/abs/2002.05202. arXiv:2002.05202 [cs]

  43. [51]

    Megatron- LM : Training Multi - Billion Parameter Language Models Using Model Parallelism , March 2020

    Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B. Megatron- LM : Training Multi - Billion Parameter Language Models Using Model Parallelism , March 2020. URL http://arxiv.org/abs/1909.08053. arXiv:1909.08053 [cs]

  44. [52]

    Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. Learning to summarize with human feedback. Advances in neural information processing systems, 33: 0 3008--3021, 2020

  45. [53]

    RoFormer : Enhanced transformer with Rotary Position Embedding

    Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. RoFormer : Enhanced transformer with Rotary Position Embedding . Neurocomputing, 568: 0 127063, February 2024. ISSN 0925-2312. doi:10.1016/j.neucom.2023.127063. URL https://www.sciencedirect.com/science/article/pii/S09252...

  46. [54]

    Llama: Open and efficient foundation language models

    Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi \`e re, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023 a

  47. [55]

    Llama 2: Open foundation and fine-tuned chat models

    Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023 b

  48. [56]

    N., Kaiser, ., and Polosukhin, I

    Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, ., and Polosukhin, I. Attention is all you need. volume 30, 2017

  49. [57]

    Auxiliary- Loss - Free Load Balancing Strategy for Mixture -of- Experts

    Wang, L., Gao, H., Zhao, C., Sun, X., and Dai, D. Auxiliary- Loss - Free Load Balancing Strategy for Mixture -of- Experts . October 2024. URL https://openreview.net/forum?id=y1iU5czYpE

  50. [58]

    Ccnet: Extracting high quality monolingual datasets from web crawl data

    Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzm \'a n, F., Joulin, A., and Grave, E. Ccnet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359, 2019

  51. [59]

    A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H

    Xiong, W., Liu, J., Molybog, I., Zhang, H., Bhargava, P., Hou, R., Martin, L., Rungta, R., Sankararaman, K. A., Oguz, B., Khabsa, M., Fang, H., Mehdad, Y., Narang, S., Malik, K., Fan, A., Bhosale, S., Edunov, S., Lewis, M., Wang, S., and Ma, H. Effective Long - Context Scaling...

  52. [60]

    and Sennrich, R

    Zhang, B. and Sennrich, R. Root mean square layer normalization. In Proceedings of the 33rd International Conference on Neural Information Processing Systems , number 1110, pp.\ 12381--12392. Curran Associates Inc., Red Hook, NY, USA, December 2019

  53. [61]

    PyTorch FSDP : Experiences on Scaling Fully Sharded Data Parallel , September 2023

    Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S. PyTorch FSDP : Experiences on Scaling Fully Sharded Data Parallel...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.