Pith. sign in

REVIEW 2 cited by

LoCo: Low-Bit Communication Adaptor for Large-scale Model Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.04480 v2 pith:5H2S4XFJ submitted 2024-07-05 cs.LG math.OC

classification cs.LGmath.OC
keywords compressionlikelocotrainingcommunicationadamgradientlarge-scale
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To efficiently train large-scale models, low-bit gradient communication compresses full-precision gradients on local GPU nodes into low-precision ones for higher gradient synchronization efficiency among GPU nodes. However, it often degrades training quality due to compression information loss. To address this, we propose the Low-bit Communication Adaptor (LoCo), which compensates gradients on local GPU nodes before compression, ensuring efficient synchronization without compromising training quality. Specifically, LoCo designs a moving average of historical compensation errors to stably estimate concurrent compression error and then adopts it to compensate for the concurrent gradient compression, yielding a less lossless compression. This mechanism allows it to be compatible with general optimizers like Adam and sharding strategies like FSDP. Theoretical analysis shows that integrating LoCo into full-precision optimizers like Adam and SGD does not impair their convergence speed on nonconvex problems. Experimental results show that across large-scale model training frameworks like Megatron-LM and PyTorch's FSDP, LoCo significantly improves communication efficiency, e.g., improving Adam's training speed by 14% to 40% without performance degradation on large language models like LLAMAs and MoE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

    cs.DC 2026-07 conditional novelty 6.0 of 10

    Transforming gradients into K-FAC-based coordinates before FP8 quantization reduces communication error and improves downstream task preservation over Euclidean FP8, with a 7.6% end-to-end speedup on 64 GH200 GPUs.

  2. Memory-Efficient 4-bit Preconditioned Stochastic Optimization

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A 4-bit Shampoo optimizer that quantizes Cholesky factors and adds error feedback matches 32-bit Shampoo's accuracy at a fraction of the memory.

Pith tools