REVIEW 24 cited by
The rising costs of training frontier AI models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The costs of training frontier AI models have grown dramatically in recent years, but there is limited public data on the magnitude and growth of these expenses. This paper develops a detailed cost model to address this gap, estimating training costs using three approaches that account for hardware, energy, cloud rental, and staff expenses. The analysis reveals that the amortized cost to train the most compute-intensive models has grown precipitously at a rate of 2.4x per year since 2016 (90% CI: 2.0x to 2.9x). For key frontier models, such as GPT-4 and Gemini, the most significant expenses are AI accelerator chips and staff costs, each costing tens of millions of dollars. Other notable costs include server components (15-22%), cluster-level interconnect (9-13%), and energy consumption (2-6%). If the trend of growing development costs continues, the largest training runs will cost more than a billion dollars by 2027, meaning that only the most well-funded organizations will be able to finance frontier AI models.
Forward citations
Cited by 24 Pith papers
-
Small-Scale Experiments: Are We There Yet?
With about 256 hyperparameter configurations per scale, scaling laws emerge at 4M parameters, and the apparent small-scale unreliability is largely a hyperparameter-tuning artifact.
-
Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms
A new Bangla–English benchmark shows vision-language models can cluster pages of shuffled government-form packets but cannot reliably reconstruct their original page order.
-
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
BaFCo introduces the first fine-grained Bangla form benchmark (26 entity types, relationships) and demonstrates that current MLLMs struggle especially with granular layout localization.
-
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
TrainVerify verifies distributed LLM training execution plans against logical model definitions using symbolic dataflow graphs, shape reduction, and staged SMT solving, scaling to 671B-parameter models.
-
zkComposer: Decomposing Proof Construction to Scale zkML
zkComposer decomposes monolithic zkML proofs into parallel sub-proofs linked by shared boundary commitments, yielding up to 6.84× lower prover time on GPT-2 without new cryptographic primitives.
-
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
Page-granular Flex-Mem and switchable Flex-NPU cut TrustZone LLM TTFT by ~10× vs a CMA strawman and ~2.4× vs a pipelined secure-NPU strawman on RK3588.
-
Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
Per-question database-style plan search over multi-LLM DAGs improves QA quality under budgets by ~58% (MMLU-Pro) and ~41% (SimpleQA) versus reimplemented baselines.
-
Optimal Resource Allocation for ML Model Training and Deployment under Concept Drift
Optimal training uses a single front-loaded burst when concept durations are DMRL, and back-loading when they are IMRL; deployment schedules are treated as quasi-convex optimization problems.
-
Unraveling Syntax: Language Modeling and the Substructure of Grammars
Language-modeling loss decomposes linearly over the sub-grammars of a probabilistic context-free grammar, and models learn these sub-grammars in parallel rather than in stages.
-
Predicting LLM Reasoning Performance with Small Proxy Model
rBridge uses a small proxy model's confidence-weighted likelihood of a frontier model's reasoning traces to predict and rank large-model reasoning performance across scales.
-
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
A skill-graph random-walk model gives closed-form accuracy-versus-compute formulas for four reasoning strategies and connects them to training scaling.
-
ZenFlow: Enabling Stall-Free Offloading Training via Asynchronous Updates
ZenFlow is an offloading framework that updates important gradients on the GPU and asynchronously accumulates unimportant ones on the CPU, achieving up to 5x speedups with comparable accuracy.
-
Position: The Most Expensive Part of an LLM should be its Training Data
Even at conservative wages, recreating LLM training data from scratch would cost 10 to 1000 times more than the compute and energy used to train the models.
-
Projected Compression: Trainable Projection for Efficient Transformer Compression
Projected Compression trains projections over frozen base model weights to produce a smaller standard transformer, outperforming hard pruning with retraining on high-token models.
-
GPTFootprint: Increasing Consumer Awareness of the Environmental Impacts of LLMs
An eco-feedback browser extension for ChatGPT raises user awareness of energy and water use, but a nine-participant study finds limited effect on query frequency.
-
Multilingual Test-Time Scaling via Initial Thought Transfer
MITT, a prefix-tuning method for multilingual test-time scaling, is evaluated on questions whose English reasoning was used for training, confounding the reported gains.
-
StruM: Structured Mixed Precision for Efficient Deep Learning Hardware Codesign
Block-wise structured mixed precision quantizes half of each weight block to low precision with under 1% ImageNet accuracy loss and powers a shifter-based accelerator PE.
-
Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability
Frontier AI access is a revocable national-security dependency, and realistic sovereignty lies in inference, deployment, and fallback capacity rather than in training frontier models.
-
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
AI replaces a human role whenever its risk-adjusted cost is lower; the paper packages this comparison with hierarchy and risk to derive conditional organizational predictions.
-
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
The paper's finding that training efficiency declines monotonically with token count is guaranteed by its efficiency metric, which divides by token count and power consumption.
-
Solving the compute crisis with physics-based ASICs
A coalition of academic and industry researchers argues that chips exploiting natural physical dynamics, rather than enforcing digital abstractions, could dramatically cut AI computing costs.
-
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
This survey classifies distributed DNN training simulators into analytical, profiling-based, and execution-driven categories, and compares them alongside TCO and carbon-emission models.
-
Bayesian Inverse Physics for Neuro-Symbolic Robot Learning
A position paper arguing that hybrid neuro-symbolic architectures combining physics, Bayesian inference, and program synthesis are essential for general-purpose robot learning.
-
The Internet of Large Language Models: An Orchestration Framework for LLM Training and Knowledge Exchange Toward Artificial General Intelligence
The paper proposes the Internet of LLM framework for model sharing, unified environments, agent-path optimization, and compute-sharing incentives, but presents no implementation or empirical evidence that it works.
Discussion (0). Continue with ORCID to comment.