Pith. sign in

REVIEW

A Diffusion Approximation Theory of Momentum SGD in Nonconvex Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.05155 v5 pith:N2HG5O3Y submitted 2018-02-14 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords momentummsgdnonconvexoptimizationannealingconvergencedeepdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Momentum Stochastic Gradient Descent (MSGD) algorithm has been widely applied to many nonconvex optimization problems in machine learning, e.g., training deep neural networks, variational Bayesian inference, and etc. Despite its empirical success, there is still a lack of theoretical understanding of convergence properties of MSGD. To fill this gap, we propose to analyze the algorithmic behavior of MSGD by diffusion approximations for nonconvex optimization problems with strict saddle points and isolated local optima. Our study shows that the momentum helps escape from saddle points, but hurts the convergence within the neighborhood of optima (if without the step size annealing or momentum annealing). Our theoretical discovery partially corroborates the empirical success of MSGD in training deep neural networks.

Discussion (0). Continue with ORCID to comment.

Pith tools