REVIEW 12 cited by
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function $f_\theta$ (which maps input vectors to output vectors) follows the kernel gradient of the functional cost (which is convex, in contrast to the parameter cost) w.r.t. a new kernel: the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and it stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We prove the positive-definiteness of the limiting NTK when the data is supported on the sphere and the non-linearity is non-polynomial. We then focus on the setting of least-squares regression and show that in the infinite-width limit, the network function $f_\theta$ follows a linear differential equation during training. The convergence is fastest along the largest kernel principal components of the input data with respect to the NTK, hence suggesting a theoretical motivation for early stopping. Finally we study the NTK numerically, observe its behavior for wide networks, and compare it to the infinite-width limit.
Forward citations
Cited by 12 Pith papers
-
Neural Spectral Bias and Conformal Correlators I: Introduction and Applications
Simple feed-forward neural networks trained on crossing symmetry plus a single anchor value reproduce CFT correlators to percent-level accuracy, and the authors conjecture this works because physical correlators are t...
-
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.
-
Coupled by Design: Computing Kerr-Newman Quasinormal Modes with a Hybrid SpectralPINN Solver
A hybrid spectral/PINN solver produces the first systematic public dataset of Kerr-Newman quasinormal-mode frequencies across the full sub-extremal parameter space, including both gravitational- and vector-led branches.
-
Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations
A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.
-
Neural Networks Reveal a Universal Bias in Conformal Correlators
Simple neural networks trained on crossing symmetry and one anchor point reproduce conformal correlators to within a few percent across many CFTs.
-
Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.
-
Quantitative Understanding of PDF Fits and their Uncertainties
After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.
-
Pre-Strings Lectures on Artificial Intelligence
Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.
-
A physics-informed neural network approach to the point defect model for electrochemical oxide film growth
A hybrid PINN anchored by one FEM data point reproduces point-defect-model film thicknesses to about 1% error, while the pure PINN overpredicts by 2,400-5,700%.
-
Is data-efficient learning feasible with quantum models?
Quantum kernels can beat an untuned classical kernel on datasets whose labels the authors deliberately construct from the quantum kernel's own spectrum, showing data efficiency by construction.
-
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.
-
Statistical Machine Learning for Astronomy -- A Textbook
A systematic Bayesian-first textbook that derives classical and modern machine learning methods for astronomy from probability theory, explicitly without new research results.
Discussion (0). Sign in to comment.