Pith. sign in

REVIEW 2 cited by

Communication Efficient Distributed Training with Distributed Lion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00438 v1 pith:Z6ASNFH3 submitted 2024-03-30 cs.DC cs.AIcs.LGmath.OCstat.ML

classification cs.DCcs.AIcs.LGmath.OCstat.ML
keywords liondistributedtrainingcommunicationadamwdemonstrateefficientgradients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Lion optimizer has been a promising competitor with the AdamW for training large AI models, with advantages on memory, computation, and sample efficiency. In this paper, we introduce Distributed Lion, an innovative adaptation of Lion for distributed training environments. Leveraging the sign operator in Lion, our Distributed Lion only requires communicating binary or lower-precision vectors between workers to the center server, significantly reducing the communication cost. Our theoretical analysis confirms Distributed Lion's convergence properties. Empirical results demonstrate its robustness across a range of tasks, worker counts, and batch sizes, on both vision and language problems. Notably, Distributed Lion attains comparable performance to standard Lion or AdamW optimizers applied on aggregated gradients, but with significantly reduced communication bandwidth. This feature is particularly advantageous for training large models. In addition, we also demonstrate that Distributed Lion presents a more favorable performance-bandwidth balance compared to existing efficient distributed methods such as deep gradient compression and ternary gradients.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise

    math.OC 2025-06 reject novelty 6.0 of 10

    Lion and Muon with weight decay are shown to be instances of one stochastic Frank-Wolfe algorithm, and clipped and variance-reduced variants get the first high-probability convergence rates for nonconvex Frank-Wolfe u...

  2. Distributed Sign Momentum with Local Steps for Training Transformers

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Distributed sign momentum with local steps matches or beats SlowMo on GPT-2 pretraining at 12x-36x lower communication and has a stated O(1/T^{1/4}) convergence rate, though the printed theorem has algebra inconsistencies.

Pith tools