Pith. sign in

REVIEW 3 cited by

Understanding Edge-of-Stability Training Dynamics with a Minimalist Example

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03294 v2 pith:SAGL45HC submitted 2022-10-07 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords sharpnesstrainingdynamicsbehaviordespiteedge-of-stabilityexamplelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/\eta$ (where $\eta$ is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below $2/\eta$. While many other well-understood nonconvex objectives such as matrix factorization or two-layer networks can also converge despite large sharpness, there is often a larger gap between sharpness of the endpoint and $2/\eta$. In this paper, we study EoS phenomenon by constructing a simple function that has the same behavior. We give rigorous analysis for its training dynamics in a large local region and explain why the final converging point has sharpness close to $2/\eta$. Globally we observe that the training dynamics for our example has an interesting bifurcating behavior, which was also observed in the training of neural nets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials

    cs.LG 2026-02 conditional novelty 7.0 of 10

    Stable oscillations of GD near sharp minima are characterized by a multivariate derivative condition, and SGD stability in expectation is governed by a worst-case batch.

  2. From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Logistic regression with gradient descent and infinite step size is the batch perceptron, and a normalized version achieves an n times better iteration complexity.

  3. Criteria and Bias of Parameterized Linear Regression under Edge of Stability Regime

    math.OC 2024-12 conditional novelty 6.0 of 10

    Under specific conditions, gradient descent converges in the unstable edge-of-stability regime for a quadratic loss on a depth-2 diagonal linear network, with a bias bound depending on step size and initialization.

Pith tools