REVIEW 36 cited by
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function $f_\theta$ (which maps input vectors to output vectors) follows the kernel gradient of the functional cost (which is convex, in contrast to the parameter cost) w.r.t. a new kernel: the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and it stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We prove the positive-definiteness of the limiting NTK when the data is supported on the sphere and the non-linearity is non-polynomial. We then focus on the setting of least-squares regression and show that in the infinite-width limit, the network function $f_\theta$ follows a linear differential equation during training. The convergence is fastest along the largest kernel principal components of the input data with respect to the NTK, hence suggesting a theoretical motivation for early stopping. Finally we study the NTK numerically, observe its behavior for wide networks, and compare it to the infinite-width limit.
Forward citations
Cited by 36 Pith papers
-
Neural Spectral Bias and Conformal Correlators I: Introduction and Applications
Neural networks optimized solely on crossing symmetry reconstruct CFT correlators from minimal input data to few-percent accuracy across generalized free fields, minimal models, Ising, N=4 SYM, and AdS diagrams.
-
Editing Models with Task Arithmetic
Task vectors from weight differences allow arithmetic operations to edit pre-trained models, improving multiple tasks simultaneously and enabling analogical inference on unseen tasks.
-
Channel Location Constrains the Auditability of Subliminal Learning
Auditability of subliminal learning is constrained by channel location, with initialization-dependent body channels allowing pre-training screens while vocabulary geometry and conditional body channels evade them.
-
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
Force-aware NTKs and chunked acquisition enable scalable, robust active learning for MLIPs, achieving lowest energy and force errors on OC20 and remaining competitive on other benchmarks.
-
Kernel-based guarantees for nonlinear parametric models in Bayesian optimization
A kernel framework over parameter space yields confidence bounds for regularized nonlinear models on adaptive data, supporting convergence analysis in Bayesian optimization.
-
Criticality and Saturation in Orthogonal Neural Networks
Derives layer-wise recursions for finite-width tensors under orthogonal initialization that reproduce the observed large-depth stability of nonlinear networks.
-
Dimensional Criticality at Grokking Across MLPs and Transformers
Effective cascade dimension D(t) crosses D=1 at the grokking transition in MLPs and Transformers, with opposite directions for modular addition versus XOR, consistent with attraction to a shared critical manifold.
-
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.
-
Coupled by Design: Computing Kerr-Newman Quasinormal Modes with a Hybrid SpectralPINN Solver
A hybrid spectral/PINN solver produces the first systematic public dataset of Kerr-Newman quasinormal-mode frequencies across the full sub-extremal parameter space, including both gravitational- and vector-led branches.
-
Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations
A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.
-
Learning from almost nothing: How neural networks survive heavy input corruption
Infinite-width MLPs implement a nearest-class-mean prototype classifier as their leading-order decision rule under heavy attribute noise, explaining observed robustness in experiments.
-
Physics-Informed Neural Networks with Attention Feature Expansion for Monge-Amp\`ere Equations
PINN-AFE uses multi-head attention and input convex networks to solve Monge-Ampère equations with claimed accuracy, efficiency, and extensions to image enhancement and medical registration.
-
Does Weight Decay Enhance Training Stability?
Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding m...
-
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
Force-aware Neural Tangent Kernels combined with chunked acquisition provide scalable and distribution-robust active learning for MLIPs, outperforming baselines on OC20 and remaining competitive on other benchmarks.
-
Neural Spectral Bias and Conformal Correlators I: Introduction and Applications
Simple feed-forward neural networks trained on crossing symmetry plus a single anchor value reproduce CFT correlators to percent-level accuracy, and the authors conjecture this works because physical correlators are t...
-
Neural Networks Reveal a Universal Bias in Conformal Correlators
Simple neural networks trained on crossing symmetry and one anchor point reproduce conformal correlators to within a few percent across many CFTs.
-
Neural Networks Reveal a Universal Bias in Conformal Correlators
Neural networks trained on crossing symmetry accurately reconstruct conformal correlators from minimal inputs due to alignment between their spectral bias and CFT smoothness.
-
Grokking as Dimensional Phase Transition in Neural Networks
Grokking occurs as the effective dimensionality of the gradient field transitions from sub-diffusive to super-diffusive at the onset of generalization, exhibiting self-organized criticality.
-
Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues
A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.
-
Quantitative Understanding of PDF Fits and their Uncertainties
After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.
-
Sharpness-Aware Minimization for Efficiently Improving Generalization
SAM solves a min-max problem to locate flat low-loss regions, improving generalization on CIFAR, ImageNet and label-noise tasks.
-
Pre-Strings Lectures on Artificial Intelligence
Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.
-
Theory of learning of high-dimensional controlled non-linear dynamical systems (I): models and methods
Introduces models for neural ODEs trained with online SGD and derives their high-dimensional learning curves via dynamical mean field theory.
-
Bayesian Inference with Shaped Deep Non-linear MLPs
In the LP/N = Θ(1) regime, Bayesian predictive posteriors for deep MLPs equal those of data-dependent kernels to first order, with a criterion identifying data processes that benefit from larger effective depth.
-
Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization
Periodic outer-momentum restarts in two-phase optimizers exploit phase cancellation in a linearized NTK model to widen stable learning-rate and momentum ranges in language-model pretraining.
-
The Thermodynamic Costs of Simple Linear Regression
Thermodynamic lower bounds are approximated for exact and SGD linear regression, producing energy-aware scaling laws for optimal training dataset size given a target generalization error.
-
A physics-informed neural network approach to the point defect model for electrochemical oxide film growth
A hybrid PINN anchored by one FEM data point reproduces point-defect-model film thicknesses to about 1% error, while the pure PINN overpredicts by 2,400-5,700%.
-
Is data-efficient learning feasible with quantum models?
Quantum kernels can beat an untuned classical kernel on datasets whose labels the authors deliberately construct from the quantum kernel's own spectrum, showing data efficiency by construction.
-
Approximating Simple ReLU Networks based on Spectral Decomposition of Fisher Information
For random 2-layer ReLU networks the dominant eigenspaces of the Fisher information matrix are spanned by spherical harmonics of degree ≤2 and capture 97.7% of the trace independently of parameter count.
-
Integrating Out, Twice:The Open-System Case That Neural-Network Ensemble Theory Is Missing
Neural-network ensembles match closed Gaussian systems but lack the open-system non-Hermitian generator and continuous spectrum required by nuclear optical models, yielding a structural negative on applicability.
-
Neural Networks, Dispersion Relations and the Thermal Bootstrap
A neural-network approach with dispersion relations handles infinite OPE towers in thermal conformal correlators without positivity.
-
On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations
The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.
-
Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation
Batch gradient descent achieves linear convergence to zero MSE with high probability for sufficiently wide shallow NNs with non-affine piecewise affine activations and distinct inputs.
-
Lectures on Semiclassical Methods for Composite Operators
Lecture notes develop semiclassical methods to compute large-n scaling dimensions of composite operators in CFTs, recovering known results in free theory and deriving one-loop corrections at the Wilson-Fisher fixed point.
-
Some Inverse Problems in Particle Physics
Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.
-
Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics
A comprehensive review of deep learning techniques for computational mechanics, including LSTM for constitutive modeling, PINNs for PDE solving, optimizers, and kernel methods.
Discussion (0). Sign in to comment.