Pith. sign in

REVIEW 24 cited by

The rising costs of training frontier AI models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.21015 v2 pith:O7JAHY5B submitted 2024-05-31 cs.CY

classification cs.CY
keywords costsmodelsfrontiertrainingcostexpensesdollarsenergy
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The costs of training frontier AI models have grown dramatically in recent years, but there is limited public data on the magnitude and growth of these expenses. This paper develops a detailed cost model to address this gap, estimating training costs using three approaches that account for hardware, energy, cloud rental, and staff expenses. The analysis reveals that the amortized cost to train the most compute-intensive models has grown precipitously at a rate of 2.4x per year since 2016 (90% CI: 2.0x to 2.9x). For key frontier models, such as GPT-4 and Gemini, the most significant expenses are AI accelerator chips and staff costs, each costing tens of millions of dollars. Other notable costs include server components (15-22%), cluster-level interconnect (9-13%), and energy consumption (2-6%). If the trend of growing development costs continues, the largest training runs will cost more than a billion dollars by 2027, meaning that only the most well-funded organizations will be able to finance frontier AI models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Small-Scale Experiments: Are We There Yet?

    cs.LG 2026-08 conditional novelty 7.0 of 10

    With about 256 hyperparameter configurations per scale, scaling laws emerge at 4M parameters, and the apparent small-scale unreliability is largely a hyperparameter-tuning artifact.

  2. Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms

    cs.CL 2026-07 conditional novelty 7.0 of 10

    A new Bangla–English benchmark shows vision-language models can cluster pages of shuffled government-form packets but cannot reliably reconstruct their original page order.

  3. BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension

    cs.CL 2026-07 accept novelty 7.0 of 10

    BaFCo introduces the first fine-grained Bangla form benchmark (26 entity types, relationships) and demonstrates that current MLLMs struggle especially with granular layout localization.

  4. TrainVerify: Equivalence-Based Verification for Distributed LLM Training

    cs.DC 2025-06 reject novelty 7.0 of 10

    TrainVerify verifies distributed LLM training execution plans against logical model definitions using symbolic dataflow graphs, shape reduction, and staged SMT solving, scaling to 671B-parameter models.

  5. zkComposer: Decomposing Proof Construction to Scale zkML

    cs.CR 2026-07 accept novelty 6.5 of 10

    zkComposer decomposes monolithic zkML proofs into parallel sub-proofs linked by shared boundary commitments, yielding up to 6.84× lower prover time on GPT-2 without new cryptographic primitives.

  6. FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation

    cs.CR 2026-03 unverdicted novelty 6.5 of 10

    Page-granular Flex-Mem and switchable Flex-NPU cut TrustZone LLM TTFT by ~10× vs a CMA strawman and ~2.4× vs a pipelined secure-NPU strawman on RK3588.

  7. Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

    cs.AI 2026-06 conditional novelty 6.0 of 10

    Per-question database-style plan search over multi-LLM DAGs improves QA quality under budgets by ~58% (MMLU-Pro) and ~41% (SimpleQA) versus reimplemented baselines.

  8. Optimal Resource Allocation for ML Model Training and Deployment under Concept Drift

    cs.LG 2025-12 reject novelty 6.0 of 10

    Optimal training uses a single front-loaded burst when concept durations are DMRL, and back-loading when they are IMRL; deployment schedules are treated as quasi-convex optimization problems.

  9. Unraveling Syntax: Language Modeling and the Substructure of Grammars

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Language-modeling loss decomposes linearly over the sub-grammars of a probabilistic context-free grammar, and models learn these sub-grammars in parallel rather than in stages.

  10. Predicting LLM Reasoning Performance with Small Proxy Model

    cs.LG 2025-09 conditional novelty 6.0 of 10

    rBridge uses a small proxy model's confidence-weighted likelihood of a frontier model's reasoning traces to predict and rank large-model reasoning performance across scales.

  11. A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.

  12. ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates

    cs.DC 2025-05 conditional novelty 6.0 of 10

    ZenFlow is an offloading framework that updates important gradients on the GPU and asynchronously accumulates unimportant ones on the CPU, achieving up to 5x speedups with comparable accuracy.

  13. Position: The Most Expensive Part of an LLM should be its Training Data

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Even at conservative wages, recreating LLM training data from scratch would cost 10 to 1000 times more than the compute and energy used to train the models.

  14. Projected Compression: Trainable Projection for Efficient Transformer Compression

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Projected Compression trains projections over frozen base model weights to produce a smaller standard transformer, outperforming hard pruning with retraining on high-token models.

  15. GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs

    cs.HC 2025-05 conditional novelty 5.0 of 10

    An eco-feedback browser extension for ChatGPT raises user awareness of energy and water use, but a nine-participant study finds limited effect on query frequency.

  16. Multilingual Test-Time Scaling via Initial Thought Transfer

    cs.CL 2025-05 reject novelty 5.0 of 10

    MITT, a prefix-tuning method for multilingual test-time scaling, is evaluated on questions whose English reasoning was used for training, confounding the reported gains.

  17. StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign

    cs.AR 2025-01 conditional novelty 5.0 of 10

    Block-wise structured mixed precision quantizes half of each weight block to low precision with under 1% ImageNet accuracy loss and powers a shifter-based accelerator PE.

  18. Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability

    cs.AI 2026-08 conditional novelty 4.0 of 10

    Frontier AI access is a revocable national-security dependency, and realistic sovereignty lies in inference, deployment, and fallback capacity rather than in training frontier models.

  19. The Human-AI Substitution Principle: When will you be replaced by AI in your organization?

    cs.AI 2026-07 conditional novelty 4.0 of 10

    AI replaces a human role whenever its risk-adjusted cost is lower; the paper packages this comparison with hierarchy and risk to derive conditional organizational predictions.

  20. Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

    cs.LG 2026-01 reject novelty 4.0 of 10

    The paper's finding that training efficiency declines monotonically with token count is guaranteed by its efficiency metric, which divides by token count and power consumption.

  21. Solving the compute crisis with physics-based ASICs

    cs.ET 2025-07 unverdicted novelty 4.0 of 10

    A coalition of academic and industry researchers argues that chips exploiting natural physical dynamics, rather than enforcing digital abstractions, could dramatically cut AI computing costs.

  22. A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO

    cs.DC 2025-06 conditional novelty 2.0 of 10

    This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.

  23. Bayesian Inverse Physics for Neuro-Symbolic Robot Learning

    cs.RO 2025-06 conditional novelty 2.0 of 10

    A position paper arguing that hybrid neuro-symbolic architectures combining physics, Bayesian inference, and program synthesis are essential for general-purpose robot learning.

  24. The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence

    cs.AI 2025-01 reject novelty 2.0 of 10

    The paper proposes the Internet of LLM framework for model sharing, unified environments, agent-path optimization, and compute-sharing incentives, but presents no implementation or empirical evidence that it works.

Pith tools