Pith. sign in

REVIEW

SGD with a Constant Large Learning Rate Can Converge to Local Maxima

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.11774 v4 pith:NI3RNNCM submitted 2021-07-25 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords localmaximaconstructconvergeslandscapeslearningoftenprevious
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous works on stochastic gradient descent (SGD) often focus on its success. In this work, we construct worst-case optimization problems illustrating that, when not in the regimes that the previous works often assume, SGD can exhibit many strange and potentially undesirable behaviors. Specifically, we construct landscapes and data distributions such that (1) SGD converges to local maxima, (2) SGD escapes saddle points arbitrarily slowly, (3) SGD prefers sharp minima over flat ones, and (4) AMSGrad converges to local maxima. We also realize results in a minimal neural network-like example. Our results highlight the importance of simultaneously analyzing the minibatch sampling, discrete-time updates rules, and realistic landscapes to understand the role of SGD in deep learning.

Discussion (0). Continue with ORCID to comment.

Pith tools