Pith. sign in

REVIEW 3 cited by

ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.18554 v1 pith:ZHJCZQAB submitted 2025-02-25 cs.DC

classification cs.DC
keywords collectivecollectivescommunicationlossyerror-boundedsignificantlyzcclcompression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the ever-increasing computing power of supercomputers and the growing scale of scientific applications, the efficiency of MPI collective communication turns out to be a critical bottleneck in large-scale distributed and parallel processing. The large message size in MPI collectives is particularly concerning because it can significantly degrade overall parallel performance. To address this issue, prior research simply applies off-the-shelf fixed-rate lossy compressors in the MPI collectives, leading to suboptimal performance, limited generalizability, and unbounded errors. In this paper, we propose a novel solution, called ZCCL, which leverages error-bounded lossy compression to significantly reduce the message size, resulting in a substantial reduction in communication costs. The key contributions are three-fold. (1) We develop two general, optimized lossy-compression-based frameworks for both types of MPI collectives (collective data movement as well as collective computation), based on their particular characteristics. Our framework not only reduces communication costs but also preserves data accuracy. (2) We customize fZ-light, an ultra-fast error-bounded lossy compressor, to meet the specific needs of collective communication. (3) We integrate ZCCL into multiple collectives, such as Allgather, Allreduce, Scatter, and Broadcast, and perform a comprehensive evaluation based on real-world scientific application datasets. Experiments show that our solution outperforms the original MPI collectives as well as multiple baselines by 1.9--8.9X.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FSZ: Breaking the Prediction-Throughput Trade-off in GPU Lossy Compression

    cs.DC 2026-07 conditional novelty 6.0 of 10

    FSZ is a single-kernel GPU lossy compressor that scores four prediction variants in one data pass, beating cuSZp-O by up to 2.92x in ratio while running at 676/785 GB/s.

  2. MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?

    cs.LG 2025-09 reject novelty 5.0 of 10

    In a 26-layer MoE model, injecting Gaussian weight errors into middle-layer experts hurts math accuracy most, while deep-layer errors can sometimes improve instruction compliance.

  3. Enhanced Sensitivity and Noise Resilience in Two-Qubit Quantum Magnetometers

    quant-ph 2025-08 unverdicted novelty 4.0 of 10

    The abstract claims a novel two-qubit magnetometer with derived sensitivity and noise measures, but the full text is an unrelated GPU all-reduce paper, so the claim is unsupported by the submission.

Pith tools