Pici-backed adaptive sparse collectives on GPUs deliver up to 5.25×/2.5×/2.66× speedups over dense NCCL for all-gather/reduce-scatter/all-reduce at 99% sparsity.
A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks,
1 Pith paper cite this work, alongside 163 external citations. Polarity classification is still indexing.
1
Pith paper citing it
163
external citations · external index
fields
cs.DC 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms
Pici-backed adaptive sparse collectives on GPUs deliver up to 5.25×/2.5×/2.66× speedups over dense NCCL for all-gather/reduce-scatter/all-reduce at 99% sparsity.