Pith. sign in

Towards better generalization: Weight decay induces low-rank bias for neural networks

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.LG 2

years

2026 2

representative citing papers

Does Weight Decay Enhance Training Stability?

cs.LG · 2026-05-15 · conditional · novelty 6.0

Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.

citing papers explorer

Showing 2 of 2 citing papers.

  • Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction cs.LG · 2026-06-04 · unverdicted · none · ref 61

    Deep linear network theory derives logarithmic decay for cross-entropy loss under gap-growth conditions versus polynomial closure for Schatten-regularized structural energy under late-time KL tails, separating fitting from simplification; conditional reductions extend this to ReLU MLPs with fixed ac

  • Does Weight Decay Enhance Training Stability? cs.LG · 2026-05-15 · conditional · none · ref 9

    Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding more stable NTK dynamics.