CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling

Dong Xu (1) , Han Meng (1) , Xinyu Chen (2) , Dengcheng Zhu (3) , Wei Tang (3) , Fei Liu (3) , Liguang Xie (3) , Wu Xiang (3)

show 9 more authors

Rui Shi (3) Yue Li (3) Henry Hu (3) Hui Zhang (3) Jianping Jiang (4) Dong Li (1) ((1) UC Merced (2) Zhejinag University (3) Bytedance (4) Xconn-tech)

Authors on Pith no claims yet

classification 💻 cs.DC cs.ET

keywords memorytimesacrosscollectivecommunicationnamenodespool

0 comments

read the original abstract

Large language models (LLMs) training or inference across multiple nodes introduces significant pressure on GPU memory and interconnect bandwidth. The Compute Express Link (CXL) shared memory pool offers a scalable solution by enabling memory sharing across nodes, reducing over-provisioning and improving resource utilization. We propose \name, a collective communication library, leveraging the CXL shared memory pool to support cross-node GPU operations without relying on traditional RDMA-based networking. Our design addresses the challenges on synchronization, data interleaving, and communication parallelization faced by using the CXL shared memory pool for collective communications. Evaluating on multiple nodes with a TITAN-II CXL switch and six Micron CZ120 memory cards, we show that \name achieves highly efficient collective operations across hosts, demonstrating CXL's potential for scalable, memory-centric GPU communication. Our evaluation demonstrates that \name achieves average performance improvements of 1.34$\times$ for AllGather, 1.84$\times$ for Broadcast, 1.94$\times$ for Gather, and 1.04$\times$ for Scatter, compared to the original RDMA-based implementation over 200 Gbps InfiniBand. \textcolor{dong}{In addition, the evaluation with a case of LLM training shows 1.11$\times$ speedup compared with the InfiniBand while saving production cost by $2.75\times$ in hardware.}

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
cs.DC 2026-04 unverdicted novelty 6.0

CoCoDiff achieves 3.6x average and 8.4x peak speedup for distributed DiT inference on up to 96 GPU tiles via tile-aware all-to-all, V-first scheduling, and selective V communication.
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
cs.OS 2026-04 unverdicted novelty 6.0

TierBPF uses lightweight eBPF hooks for custom page admission control in tiered memory, delivering up to 17.7% geomean and 75% peak throughput gains across 17 workloads on three systems.
Hybrid Adaptive Tuning for Tiered Memory Systems
cs.OS 2026-04 unverdicted novelty 6.0

PTMT is a lightweight framework that automates parameter tuning for memory tiering via hybrid offline database building and online customized reinforcement learning, delivering 14-30% gains over defaults and 32% over ...