ZeCO sequence parallelism for linear attention uses a pipelined All-Scan collective to cut communication volume and time, achieving near-linear scaling up to 256 GPUs with 8M-token sequences.
The experimental setup with 5 rounds of warm-up reported the average of 100 steps of the experiment, see in Table 5, Table
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
ZeCO sequence parallelism for linear attention uses a pipelined All-Scan collective to cut communication volume and time, achieving near-linear scaling up to 256 GPUs with 8M-token sequences.