REVIEW 6 cited by
Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Given a non-convex twice differentiable cost function f, we prove that the set of initial conditions so that gradient descent converges to saddle points where \nabla^2 f has at least one strictly negative eigenvalue has (Lebesgue) measure zero, even for cost functions f with non-isolated critical points, answering an open question in [Lee, Simchowitz, Jordan, Recht, COLT2016]. Moreover, this result extends to forward-invariant convex subspaces, allowing for weak (non-globally Lipschitz) smoothness assumptions. Finally, we produce an upper bound on the allowable step-size.
Forward citations
Cited by 6 Pith papers
-
Limitations of Lazy Training of Two-layers Neural Networks
For quadratic targets in d dimensions, two-layer quadratic networks achieve lower risk when fully trained than in random features or neural tangent regimes if hidden units < d.
-
Accelerated Gradient Methods for Nonconvex Optimization: Escape Trajectories From Strict Saddle Points and Convergence to Local Minima
Theoretical analysis of accelerated gradient methods showing almost-sure escape from strict saddles and linear exit times, plus a subclass achieving near-optimal convergence to local minima in convex neighborhoods of ...
-
On The Geometric Analysis of A Quartic-quadratic Optimization Problem under A Spherical Constraint
For the quartic-quadratic sphere problem, the paper characterizes all local minima in the diagonal case, proves strict-saddle properties for extreme beta, and establishes a Kurdyka-Lojasiewicz exponent of 1/4.
-
Landscape analysis for shallow neural networks: Complete classification of critical points for cubic activation and affine target functions
For cubic-activation shallow networks with affine targets, the squared-loss landscape has no local maxima; every critical point is a global minimizer, a rigid non-global local minimum, or a saddle, and zero loss is ac...
-
On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem
A Riemannian subgradient differential inclusion unifies Hessian barrier and mirror descent methods and explains their spurious stationary points as stable equilibria outside the true stationary set.
-
Extending the step-size restriction for gradient descent to avoid strict saddle points
Gradient descent almost surely avoids strict saddles for step sizes alpha < 2/L, provided the Hessian eigenvalue alpha^-1 occurs only on a measure-zero set.
Discussion (0). Continue with ORCID to comment.