Classical momentum acceleration in mini-batch SGD for quadratics is proportional to batch size up to saturation, enabling perfect parallelization under minimal noise assumptions.
Linearly convergent stochastic heavy ball method for minimizing generalization error
2 Pith papers cite this work. Polarity classification is still indexing.
abstract
In this work we establish the first linear convergence result for the stochastic heavy ball method. The method performs SGD steps with a fixed stepsize, amended by a heavy ball momentum term. In the analysis, we focus on minimizing the expected loss and not on finite-sum minimization, which is typically a much harder problem. While in the analysis we constrain ourselves to quadratic loss, the overall objective is not necessarily strongly convex.
verdicts
UNVERDICTED 2representative citing papers
Heavy-ball methods with random starts provably escape saddle points via a new state-space mapping that allows larger steps than plain gradient descent.
citing papers explorer
-
Perfect Parallelization in Mini-Batch SGD with Classical Momentum Acceleration
Classical momentum acceleration in mini-batch SGD for quadratics is proportional to batch size up to saturation, enabling perfect parallelization under minimal noise assumptions.
-
Heavy-ball Algorithms Always Escape Saddle Points
Heavy-ball methods with random starts provably escape saddle points via a new state-space mapping that allows larger steps than plain gradient descent.