Introduces Label-NTK and Residual-NTK alignments to derive tighter NTK convergence bounds that track the full eigen-spectrum and match observed training speed.
Spectrum dependent learning curves in kernel regression and wide neural networks
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Derives α^{-1/3} scaling for generalization error in online softmax classification from boundary layers in a teacher-student model.
Spectral analysis of activations and gradients provides new diagnostics that link batch size to representation geometry, early covariance tails to token efficiency, and spectral shifts to learning dynamics in decoder-only LLMs, backed by a mechanistic model.
citing papers explorer
-
Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime
Introduces Label-NTK and Residual-NTK alignments to derive tighter NTK convergence bounds that track the full eigen-spectrum and match observed training speed.
-
A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
Derives α^{-1/3} scaling for generalization error in online softmax classification from boundary layers in a teacher-student model.
-
Spectral Lens: Activation and Gradient Spectra as Diagnostics of LLM Optimization
Spectral analysis of activations and gradients provides new diagnostics that link batch size to representation geometry, early covariance tails to token efficiency, and spectral shifts to learning dynamics in decoder-only LLMs, backed by a mechanistic model.