Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

· 2016 · math.DS · arXiv 1605.00405

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

open full Pith review browse 2 citing papers arXiv PDF

abstract

Given a non-convex twice differentiable cost function f, we prove that the set of initial conditions so that gradient descent converges to saddle points where \nabla^2 f has at least one strictly negative eigenvalue has (Lebesgue) measure zero, even for cost functions f with non-isolated critical points, answering an open question in [Lee, Simchowitz, Jordan, Recht, COLT2016]. Moreover, this result extends to forward-invariant convex subspaces, allowing for weak (non-globally Lipschitz) smoothness assumptions. Finally, we produce an upper bound on the allowable step-size.

representative citing papers

Limitations of Lazy Training of Two-layers Neural Networks

stat.ML · 2019-06-21 · unverdicted · novelty 8.0

For quadratic targets in d dimensions, two-layer quadratic networks achieve lower risk when fully trained than in random features or neural tangent regimes if hidden units < d.

Accelerated Gradient Methods for Nonconvex Optimization: Escape Trajectories From Strict Saddle Points and Convergence to Local Minima

math.OC · 2023-07-13 · unverdicted · novelty 7.0

Theoretical analysis of accelerated gradient methods showing almost-sure escape from strict saddles and linear exit times, plus a subclass achieving near-optimal convergence to local minima in convex neighborhoods of nonconvex functions.

citing papers explorer

Showing 2 of 2 citing papers.

Limitations of Lazy Training of Two-layers Neural Networks stat.ML · 2019-06-21 · unverdicted · none · ref 29 · internal anchor
For quadratic targets in d dimensions, two-layer quadratic networks achieve lower risk when fully trained than in random features or neural tangent regimes if hidden units < d.
Accelerated Gradient Methods for Nonconvex Optimization: Escape Trajectories From Strict Saddle Points and Convergence to Local Minima math.OC · 2023-07-13 · unverdicted · none · ref 67 · internal anchor
Theoretical analysis of accelerated gradient methods showing almost-sure escape from strict saddles and linear exit times, plus a subclass achieving near-optimal convergence to local minima in convex neighborhoods of nonconvex functions.

Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

fields

years

verdicts

representative citing papers

citing papers explorer