REVIEW 2 cited by
The Limits and Potentials of Local SGD for Distributed Heterogeneous Learning with Intermittent Communication
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Local SGD is a popular optimization method in distributed learning, often outperforming other algorithms in practice, including mini-batch SGD. Despite this success, theoretically proving the dominance of local SGD in settings with reasonable data heterogeneity has been difficult, creating a significant gap between theory and practice. In this paper, we provide new lower bounds for local SGD under existing first-order data heterogeneity assumptions, showing that these assumptions are insufficient to prove the effectiveness of local update steps. Furthermore, under these same assumptions, we demonstrate the min-max optimality of accelerated mini-batch SGD, which fully resolves our understanding of distributed optimization for several problem classes. Our results emphasize the need for better models of data heterogeneity to understand the effectiveness of local SGD in practice. Towards this end, we consider higher-order smoothness and heterogeneity assumptions, providing new upper bounds that imply the dominance of local SGD over mini-batch SGD when data heterogeneity is low.
Forward citations
Cited by 2 Pith papers
-
Decoupled SGDA for Games with Intermittent Strategy Communication
Decoupled SGDA achieves O(1/(1-4κ_c) log(1/ϵ)) communication rounds in weakly coupled SCSC games, independent of the players' condition numbers.
-
Task Arithmetic Through The Lens Of One-Shot Federated Learning
Task arithmetic is exactly one-shot FedAvg with outer step size beta = lambda T, and FedNova, FedGMA, Median, and CCLIP can often improve merged model performance.
Discussion (0). Continue with ORCID to comment.