Non-Euclidean Gradient Descent Operates at the Edge of Stability

Jeremy Cohen; Michael Crawshaw; Robert Gower; Rustem Islamov

arxiv: 2603.05002 · v2 · pith:YNOJFVNMnew · submitted 2026-03-05 · 💻 cs.LG · math.OC· stat.ML

Non-Euclidean Gradient Descent Operates at the Edge of Stability

Rustem Islamov , Michael Crawshaw , Jeremy Cohen , Robert Gower This is my paper

classification 💻 cs.LG math.OCstat.ML

keywords non-euclideansharpnessdescentgeneralizedgradientstabilitybeenedge

0 comments

read the original abstract

The Edge of Stability (EoS) is a phenomenon where the sharpness (largest eigenvalue) of the Hessian approaches and then hovers near the stability threshold $2/\eta$ during gradient descent (GD) with step size $\eta$. Despite (apparently) violating classical smoothness assumptions, EoS has been widely observed in deep learning, but its theoretical foundations remain incomplete. We provide an interpretation of EoS through the lens of Directional Smoothness [Mishkin et al., 2024]. This interpretation naturally extends to non-Euclidean norms, which we use to define generalized sharpness under an arbitrary norm. Our generalized sharpness measure includes previously studied vanilla GD and preconditioned GD as special cases, as well as methods for which EoS has not been studied, such as $\ell_{\infty}$-descent, Block CD, Spectral GD, and their normalized versions. Through experiments on neural networks, we show that non-Euclidean GD with our generalized sharpness also exhibits progressive sharpening followed by oscillations around or above the threshold $2/\eta$. Practically, our framework provides a geometry-aware spectral diagnostic that can be applied across a broad class of non-Euclidean gradient methods.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Muon is Not That Special: Random or Inverted Spectra Work Just as Well
cs.LG 2026-05 unverdicted novelty 7.0

Muon succeeds by guaranteeing local step-size optimality rather than by tracking any ideal global geometry, as random-spectrum and quasi-norm variants match its performance on language models.
Does Weight Decay Enhance Training Stability?
cs.LG 2026-05 conditional novelty 6.0

Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding m...