A framework called H2 combines a unified PyTorch interface, device-direct RDMA, and automatically searched pipeline parallelism to train a 100B model on over 1,000 heterogeneous chips, with up to 16.37% higher aggregate throughput than separate homogeneous runs.
2016.{TensorFlow}: a system for{Large-Scale} machine learning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips
A framework called H2 combines a unified PyTorch interface, device-direct RDMA, and automatically searched pipeline parallelism to train a 100B model on over 1,000 heterogeneous chips, with up to 16.37% higher aggregate throughput than separate homogeneous runs.