Pith. sign in

REVIEW 1 cited by

On the convergence properties of a $K$-step averaging stochastic gradient descent algorithm for nonconvex optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.01012 v3 pith:F4H5QNAJ submitted 2017-08-03 cs.LG cs.DCstat.ML

classification cs.LGcs.DCstat.ML
keywords k-avgasgddescentgradientstochasticbetterconvergencealgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Despite their popularity, the practical performance of asynchronous stochastic gradient descent methods (ASGD) for solving large scale machine learning problems are not as good as theoretical results indicate. We adopt and analyze a synchronous K-step averaging stochastic gradient descent algorithm which we call K-AVG. We establish the convergence results of K-AVG for nonconvex objectives and explain why the K-step delay is necessary and leads to better performance than traditional parallel stochastic gradient descent which is a special case of K-AVG with $K=1$. We also show that K-AVG scales better than ASGD. Another advantage of K-AVG over ASGD is that it allows larger stepsizes. On a cluster of $128$ GPUs, K-AVG is faster than ASGD implementations and achieves better accuracies and faster convergence for \cifar dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Batch-Size Control for Low-Latency Federated Learning with Device Heterogeneity

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Wireless federated learning can cut end-to-end training time by choosing per-device batch sizes with a closed-form rule that balances convergence rounds against per-round latency.

Pith tools