Pith. sign in

REVIEW 1 cited by

Improved Quantization Strategies for Managing Heavy-tailed Gradients in Distributed Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.01798 v1 pith:H7VWLBGB submitted 2024-02-02 cs.LG cs.DC

classification cs.LGcs.DC
keywords quantizationdistributedheavy-tailedgradientgradientslearningcompressionerror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Gradient compression has surfaced as a key technique to address the challenge of communication efficiency in distributed learning. In distributed deep learning, however, it is observed that gradient distributions are heavy-tailed, with outliers significantly influencing the design of compression strategies. Existing parameter quantization methods experience performance degradation when this heavy-tailed feature is ignored. In this paper, we introduce a novel compression scheme specifically engineered for heavy-tailed gradients, which effectively combines gradient truncation with quantization. This scheme is adeptly implemented within a communication-limited distributed Stochastic Gradient Descent (SGD) framework. We consider a general family of heavy-tail gradients that follow a power-law distribution, we aim to minimize the error resulting from quantization, thereby determining optimal values for two critical parameters: the truncation threshold and the quantization density. We provide a theoretical analysis on the convergence error bound under both uniform and non-uniform quantization scenarios. Comparative experiments with other benchmarks demonstrate the effectiveness of our proposed method in managing the heavy-tailed gradients in a distributed learning environment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lion Cub: Minimizing Communication Overhead in Distributed Lion

    cs.LG 2024-11 conditional novelty 6.0 of 10

    Lion Cub compresses Lion updates with L1 quantization and sparse momentum synchronization, reducing distributed training time by up to 5.1x at similar convergence.

Pith tools