Pith. sign in

REVIEW 3 cited by

Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13093 v4 pith:IBQUSMZN submitted 2023-03-23 cs.LG math.OCphysics.data-an

classification cs.LGmath.OCphysics.data-an
keywords dynamicssaddlegradientsaddlesaroundpointsdescentprobabilistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Characterizing and understanding the dynamics of stochastic gradient descent (SGD) around saddle points remains an open problem. We first show that saddle points in neural networks can be divided into two types, among which the Type-II saddles are especially difficult to escape from because the gradient noise vanishes at the saddle. The dynamics of SGD around these saddles are thus to leading order described by a random matrix product process, and it is thus natural to study the dynamics of SGD around these saddles using the notion of probabilistic stability and the related Lyapunov exponent. Theoretically, we link the study of SGD dynamics to well-known concepts in ergodic theory, which we leverage to show that saddle points can be either attractive or repulsive for SGD, and its dynamics can be classified into four different phases, depending on the signal-to-noise ratio in the gradient close to the saddle.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zeroth-Order Optimization at the Edge of Stability

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Zeroth-order methods achieve mean-square stability when the step size satisfies a condition involving the entire Hessian spectrum, with full-batch ZO optimizers operating at the edge of stability and large steps regul...

  2. On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Stable oscillations of GD near sharp minima are characterized by a multivariate derivative condition, and SGD stability in expectation is governed by a worst-case batch.

  3. Deep Weight Factorization: Sparse Learning Through the Lens of Artificial Symmetries

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Factorizing weights into D≥2 multiplicative factors and applying L2 weight decay induces a non-convex sparse L2/D penalty, and with tailored initialization and learning rates, achieves superior sparsity-accuracy tradeoffs.

Pith tools