ZeCO sequence parallelism for linear attention uses a pipelined All-Scan collective to cut communication volume and time, achieving near-linear scaling up to 256 GPUs with 8M-token sequences.
The experimental setup with 5 rounds of warm-up and reported the average of 50 rounds of experiment, see in Table 3, Table
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
ZeCO sequence parallelism for linear attention uses a pipelined All-Scan collective to cut communication volume and time, achieving near-linear scaling up to 256 GPUs with 8M-token sequences.