Momentum SGD exhibits two distinct EoSS regimes for batch sharpness, stabilizing at 2(1-β)/η for small batches and 2(1+β)/η for large batches, aligning with linear stability thresholds.
NeurIPS Paper Checklist
4 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 4verdicts
UNVERDICTED 4roles
background 1polarities
background 1representative citing papers
Derives closed-form gradient of WS upper bound on Hessian max eigenvalue for 3-layer cross-entropy NNs and proposes HSR regularization to steer toward flat minima.
A closed-form upper bound on the maximum Hessian eigenvalue of cross-entropy loss is derived for smooth nonlinear neural networks.
A linear stability analysis introduces data coherence to explain why SGD and SAM prefer stable and simple minima in two-layer ReLU networks.
citing papers explorer
-
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
Momentum SGD exhibits two distinct EoSS regimes for batch sharpness, stabilizing at 2(1-β)/η for small batches and 2(1+β)/η for large batches, aligning with linear stability thresholds.
-
Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks
Derives closed-form gradient of WS upper bound on Hessian max eigenvalue for 3-layer cross-entropy NNs and proposes HSR regularization to steer toward flat minima.
-
Wolkowicz-Styan Upper Bound on the Hessian Eigenspectrum for Cross-Entropy Loss in Nonlinear Smooth Neural Networks
A closed-form upper bound on the maximum Hessian eigenvalue of cross-entropy loss is derived for smooth nonlinear neural networks.
-
A Unified Stability Analysis of SAM vs SGD: Role of Data Coherence and Emergence of Simplicity Bias
A linear stability analysis introduces data coherence to explain why SGD and SAM prefer stable and simple minima in two-layer ReLU networks.