Pith. sign in

REVIEW 1 cited by

Stochastic Gradient Descent with Polyak's Learning Rate

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1903.08688 v2 pith:SZMJ446V submitted 2019-03-20 math.OC

classification math.OC
keywords ratedescentgradientstochasticconstantconvexlearningpolyak
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Stochastic gradient descent (SGD) for strongly convex functions converges at the rate $\bO(1/k)$. However, achieving good results in practice requires tuning the parameters (for example the learning rate) of the algorithm. In this paper we propose a generalization of the Polyak step size, used for subgradient methods, to Stochastic gradient descent. We prove a non-asymptotic convergence at the rate $\bO(1/k)$ with a rate constant which can be better than the corresponding rate constant for optimally scheduled SGD. We demonstrate that the method is effective in practice, and on convex optimization problems and on training deep neural networks, and compare to the theoretical rate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nesterov's method with decreasing learning rate leads to accelerated stochastic gradient descent

    math.OC 2019-08 conditional novelty 6.0 of 10

    A coupled ODE system, discretized with a decreasing learning rate, yields accelerated SGD algorithms with proven optimal last-iterate rates and improved constants.

Pith tools