Pith. sign in

REVIEW 6 cited by

Gradient Descent Only Converges to Minimizers: Non-Isolated Critical Points and Invariant Regions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1605.00405 v2 pith:UWRCZMH6 submitted 2016-05-02 math.DS cs.LG

classification math.DScs.LG
keywords pointsconvergescostcriticaldescentgradientnon-isolatedallowable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given a non-convex twice differentiable cost function f, we prove that the set of initial conditions so that gradient descent converges to saddle points where \nabla^2 f has at least one strictly negative eigenvalue has (Lebesgue) measure zero, even for cost functions f with non-isolated critical points, answering an open question in [Lee, Simchowitz, Jordan, Recht, COLT2016]. Moreover, this result extends to forward-invariant convex subspaces, allowing for weak (non-globally Lipschitz) smoothness assumptions. Finally, we produce an upper bound on the allowable step-size.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Limitations of Lazy Training of Two-layers Neural Networks

    stat.ML 2019-06 unverdicted novelty 8.0 of 10

    For quadratic targets in d dimensions, two-layer quadratic networks achieve lower risk when fully trained than in random features or neural tangent regimes if hidden units < d.

  2. Accelerated Gradient Methods for Nonconvex Optimization: Escape Trajectories From Strict Saddle Points and Convergence to Local Minima

    math.OC 2023-07 unverdicted novelty 7.0 of 10

    Theoretical analysis of accelerated gradient methods showing almost-sure escape from strict saddles and linear exit times, plus a subclass achieving near-optimal convergence to local minima in convex neighborhoods of ...

  3. On The Geometric Analysis of A Quartic-quadratic Optimization Problem under A Spherical Constraint

    math.OC 2019-08 conditional novelty 7.0 of 10

    For the quartic-quadratic sphere problem, the paper characterizes all local minima in the diagonal case, proves strict-saddle properties for extreme beta, and establishes a Kurdyka-Lojasiewicz exponent of 1/4.

  4. Landscape analysis for shallow neural networks: Complete classification of critical points for cubic activation and affine target functions

    math.OC 2026-07 conditional novelty 6.0 of 10

    For cubic-activation shallow networks with affine targets, the squared-loss landscape has no local maxima; every critical point is a global minimizer, a rigid non-global local minimum, or a saddle, and zero loss is ac...

  5. On exploration of an interior mirror descent flow for stochastic nonconvex constrained problem

    math.OC 2025-07 conditional novelty 6.0 of 10

    A Riemannian subgradient differential inclusion unifies Hessian barrier and mirror descent methods and explains their spurious stationary points as stable equilibria outside the true stationary set.

  6. Extending the step-size restriction for gradient descent to avoid strict saddle points

    stat.ML 2019-08 conditional novelty 6.0 of 10

    Gradient descent almost surely avoids strict saddles for step sizes alpha < 2/L, provided the Hessian eigenvalue alpha^-1 occurs only on a measure-zero set.

Pith tools