Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
On large-batch training for deep learning: Generalization gap and sharp minima
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.LG 5roles
background 1polarities
background 1representative citing papers
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.
Even with perfectly balanced and whitened data, a deep linear network's Hessian exhibits a two-cluster spectrum whose dominant-to-bulk eigenvalue ratio grows linearly with depth.
C3PO is a foundation model for bilevel pricing optimization that trains on simulated discrete choice data and retrieves elasticity priors from literature to improve revenue KPIs under business constraints.
citing papers explorer
-
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Slingshot loss spikes are produced by low-precision arithmetic that breaks the zero-sum gradient constraint and drives exponential growth via Numerical Feature Inflation.
-
A Ridge Too Far: Correcting Over-Shrinkage via Negative Regularization
Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.
-
How Far Can Sharpness and Complexity Jointly Explain Generalization?
Function-space definitions of sharpness and complexity jointly explain more generalization variance than parameter-space versions, yet leave unexplained cases that suggest the two-factor view is incomplete.
-
Depth, Not Data: An Analysis of Hessian Spectral Bifurcation
Even with perfectly balanced and whitened data, a deep linear network's Hessian exhibits a two-cluster spectrum whose dominant-to-bulk eigenvalue ratio grows linearly with depth.
-
Causal-Aware Foundation-Model for Bilevel Optimization in Discrete Choice Settings
C3PO is a foundation model for bilevel pricing optimization that trains on simulated discrete choice data and retrieves elasticity priors from literature to improve revenue KPIs under business constraints.