REVIEW 3 cited by
Understanding Edge-of-Stability Training Dynamics with a Minimalist Example
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/\eta$ (where $\eta$ is the step size). Despite this, the loss oscillates and converges in the long run, and the sharpness at the end is just slightly below $2/\eta$. While many other well-understood nonconvex objectives such as matrix factorization or two-layer networks can also converge despite large sharpness, there is often a larger gap between sharpness of the endpoint and $2/\eta$. In this paper, we study EoS phenomenon by constructing a simple function that has the same behavior. We give rigorous analysis for its training dynamics in a large local region and explain why the final converging point has sharpness close to $2/\eta$. Globally we observe that the training dynamics for our example has an interesting bifurcating behavior, which was also observed in the training of neural nets.
Forward citations
Cited by 3 Pith papers
-
On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials
Stable oscillations of GD near sharp minima are characterized by a multivariate derivative condition, and SGD stability in expectation is governed by a worst-case batch.
-
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
Logistic regression with gradient descent and infinite step size is the batch perceptron, and a normalized version achieves an n times better iteration complexity.
-
Criteria and Bias of Parameterized Linear Regression under Edge of Stability Regime
Under specific conditions, gradient descent converges in the unstable edge-of-stability regime for a quadratic loss on a depth-2 diagonal linear network, with a bias bound depending on step size and initialization.
Discussion (0). Continue with ORCID to comment.