HCCL offloads collective communication to MTIA 300's message engines, achieving up to 940 GB/s intra-rack bandwidth and sub-6µs latency for inference.
Scalable Hierarchical Aggregation and Reduction Protocol (SHARP): A Hardware Architecture for Efficient Data Reduction,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.NI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HCCL: Collective Communication for Meta Training and Inference Accelerators
HCCL offloads collective communication to MTIA 300's message engines, achieving up to 940 GB/s intra-rack bandwidth and sub-6µs latency for inference.