Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
The implicit bias of gradient descent on separable data.Journal of Machine Learning Research, 19(70):1–57
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
In the high-dimensional regime, SGD on diagonal linear networks is approximated by an SDE and a deterministic PDE that together give an explicit non-asymptotic description of convergence to zero risk.
Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.
Uniform-based discrete diffusion models behave as associative memories that retrieve unseen data, with a dataset-size-driven memorization-to-generalization transition detectable via conditional entropy of token predictions.
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.
citing papers explorer
-
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
-
High-dimensional Limit of SGD for Diagonal Linear Networks
In the high-dimensional regime, SGD on diagonal linear networks is approximated by an SDE and a deterministic PDE that together give an explicit non-asymptotic description of convergence to zero risk.
-
Does Weight Decay Enhance Training Stability?
Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.
-
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
Uniform-based discrete diffusion models behave as associative memories that retrieve unseen data, with a dataset-size-driven memorization-to-generalization transition detectable via conditional entropy of token predictions.
-
A Ridge Too Far: Correcting Over-Shrinkage via Negative Regularization
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
-
How Far Can Sharpness and Complexity Jointly Explain Generalization?
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.