TrainMover achieves ~20s downtime for interruptions in 1024-GPU LLM training via two-phase delta-based communication setup, communication-free sandboxed warmup, and general standby design, projecting 55% reduction in wasted GPU hours.
https: //engineering.fb.com/2024/06/12/production-engineering/ maintaining-large-scale-ai-capacity-meta/
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
TrainMover: An Interruption-Resilient Runtime for ML Training
TrainMover achieves ~20s downtime for interruptions in 1024-GPU LLM training via two-phase delta-based communication setup, communication-free sandboxed warmup, and general standby design, projecting 55% reduction in wasted GPU hours.