REVIEW 3 cited by
Data movement limits to frontier model training
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We present a theoretical model of distributed training, and use it to analyze how far dense and sparse training runs can be scaled. Under our baseline assumptions, given a three month training duration, data movement bottlenecks begin to significantly lower hardware utilization for training runs exceeding about $10^{28}$ FLOP, two orders of magnitude above the largest training run to date, suggesting the arrival of fundamental barriers to scaling in three years given recent rates of growth. A training run exceeding about $10^{31}$ FLOP is infeasible even at low utilization. However, more aggressive batch size scaling and/or shorter and fatter model shapes, if achievable, have the potential to permit much larger training runs.
Forward citations
Cited by 3 Pith papers
-
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
When a cosine schedule would halve the learning rate, Seesaw cuts it by √2 and doubles the batch, matching loss curves with ~36% fewer serial steps.
-
Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment
Countries could verify compliance with international AI agreements through six redundant verification layers, provided the report's listed hardware and analysis challenges are solved.
-
Technical Options for Flexible Hardware-Enabled Guarantees
A hardware 'interlock' placed on AI accelerator network paths could provide privacy-preserving, verifiable guarantees about AI compute usage, according to a design analysis that sketches FLOP-counting and update protocols.
Discussion (0). Continue with ORCID to comment.