Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
Gradient descent maximizes the margin of homogeneous neural networks
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.
citing papers explorer
-
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
-
How Far Can Sharpness and Complexity Jointly Explain Generalization?
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.