Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.
How neural networks learn the support is an implicit regularization effect of sgd
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2representative citing papers
HNNs recover known sparse hierarchies on synthetic tasks and match or exceed dense DNNs on real datasets while using orders of magnitude fewer parameters and showing lower hyperparameter sensitivity.
citing papers explorer
-
Does Weight Decay Enhance Training Stability?
Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.
-
Compositional Sparsity as an Inductive Bias for Neural Architecture Design
HNNs recover known sparse hierarchies on synthetic tasks and match or exceed dense DNNs on real datasets while using orders of magnitude fewer parameters and showing lower hyperparameter sensitivity.