A quantized two-step all-reduce kernel reduces tensor-parallel communication overhead in LLM inference, achieving up to 3.18x faster all-reduce and 2.06x TTFT speedup on L40 GPUs.
baidu-allreduce: A C++ library demonstrating ring allreduce and ring allgather techniques
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference
A quantized two-step all-reduce kernel reduces tensor-parallel communication overhead in LLM inference, achieving up to 3.18x faster all-reduce and 2.06x TTFT speedup on L40 GPUs.