AxoNN combines 3D parallel matrix multiplication with data parallelism to reach 1.423 exaflop/s on 6,144 H100 GPUs, and reports one-pass catastrophic memorization at the 70B scale that a masked-loss technique suppresses.
MegaScale: Scaling large language model training to more than 10,000 GPUs,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
AxoNN combines 3D parallel matrix multiplication with data parallelism to reach 1.423 exaflop/s on 6,144 H100 GPUs, and reports one-pass catastrophic memorization at the 70B scale that a masked-loss technique suppresses.