Pith. sign in

REVIEW 36 cited by

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.07572 v4 pith:DPSOYAJ3 submitted 2018-06-20 cs.LG cs.NEmath.PRstat.ML

classification cs.LGcs.NEmath.PRstat.ML
keywords kerneltrainingduringinfinite-widthlimitneuralannsconvergence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function $f_\theta$ (which maps input vectors to output vectors) follows the kernel gradient of the functional cost (which is convex, in contrast to the parameter cost) w.r.t. a new kernel: the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and it stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We prove the positive-definiteness of the limiting NTK when the data is supported on the sphere and the non-linearity is non-polynomial. We then focus on the setting of least-squares regression and show that in the infinite-width limit, the network function $f_\theta$ follows a linear differential equation during training. The convergence is fastest along the largest kernel principal components of the input data with respect to the NTK, hence suggesting a theoretical motivation for early stopping. Finally we study the NTK numerically, observe its behavior for wide networks, and compare it to the infinite-width limit.

Discussion (0). Sign in to comment.

Forward citations

Cited by 36 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Neural Spectral Bias and Conformal Correlators I: Introduction and Applications

    hep-th 2026-04 unverdicted novelty 8.0 of 10

    Neural networks optimized solely on crossing symmetry reconstruct CFT correlators from minimal input data to few-percent accuracy across generalized free fields, minimal models, Ising, N=4 SYM, and AdS diagrams.

  2. Editing Models with Task Arithmetic

    cs.LG 2022-12 accept novelty 8.0 of 10

    Task vectors from weight differences allow arithmetic operations to edit pre-trained models, improving multiple tasks simultaneously and enabling analogical inference on unseen tasks.

  3. Channel Location Constrains the Auditability of Subliminal Learning

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Auditability of subliminal learning is constrained by channel location, with initialization-dependent body channels allowing pre-training screens while vocabulary geometry and conditional body channels evade them.

  4. Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Force-aware NTKs and chunked acquisition enable scalable, robust active learning for MLIPs, achieving lowest energy and force errors on OC20 and remaining competitive on other benchmarks.

  5. Kernel-based guarantees for nonlinear parametric models in Bayesian optimization

    stat.ML 2026-05 unverdicted novelty 7.0 of 10

    A kernel framework over parameter space yields confidence bounds for regularized nonlinear models on adaptive data, supporting convergence analysis in Bayesian optimization.

  6. Criticality and Saturation in Orthogonal Neural Networks

    cs.LG 2026-05 conditional novelty 7.0 of 10

    Derives layer-wise recursions for finite-width tensors under orthogonal initialization that reproduce the observed large-depth stability of nonlinear networks.

  7. Dimensional Criticality at Grokking Across MLPs and Transformers

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Effective cascade dimension D(t) crosses D=1 at the grokking transition in MLPs and Transformers, with opposite directions for modular addition versus XOR, consistent with attraction to a shared critical manifold.

  8. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  9. Coupled by Design: Computing Kerr-Newman Quasinormal Modes with a Hybrid SpectralPINN Solver

    gr-qc 2026-07 conditional novelty 6.0 of 10

    A hybrid spectral/PINN solver produces the first systematic public dataset of Kerr-Newman quasinormal-mode frequencies across the full sub-extremal parameter space, including both gravitational- and vector-led branches.

  10. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.

  11. Learning from almost nothing: How neural networks survive heavy input corruption

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Infinite-width MLPs implement a nearest-class-mean prototype classifier as their leading-order decision rule under heavy attribute noise, explaining observed robustness in experiments.

  12. Physics-Informed Neural Networks with Attention Feature Expansion for Monge-Amp\`ere Equations

    math.NA 2026-05 unverdicted novelty 6.0 of 10

    PINN-AFE uses multi-head attention and input convex networks to solve Monge-Ampère equations with claimed accuracy, efficiency, and extensions to image enhancement and medical registration.

  13. Does Weight Decay Enhance Training Stability?

    cs.LG 2026-05 conditional novelty 6.0 of 10

    Weight decay slows progressive sharpening at the edge of stability, inducing damped oscillations in CNNs and a phase transition to sub-2/η sharpness in MLPs driven by parameter-sharpness gradient alignment, yielding m...

  14. Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Force-aware Neural Tangent Kernels combined with chunked acquisition provide scalable and distribution-robust active learning for MLIPs, outperforming baselines on OC20 and remaining competitive on other benchmarks.

  15. Neural Spectral Bias and Conformal Correlators I: Introduction and Applications

    hep-th 2026-04 conditional novelty 6.0 of 10

    Simple feed-forward neural networks trained on crossing symmetry plus a single anchor value reproduce CFT correlators to percent-level accuracy, and the authors conjecture this works because physical correlators are t...

  16. Neural Networks Reveal a Universal Bias in Conformal Correlators

    hep-th 2026-04 conditional novelty 6.0 of 10

    Simple neural networks trained on crossing symmetry and one anchor point reproduce conformal correlators to within a few percent across many CFTs.

  17. Neural Networks Reveal a Universal Bias in Conformal Correlators

    hep-th 2026-04 unverdicted novelty 6.0 of 10

    Neural networks trained on crossing symmetry accurately reconstruct conformal correlators from minimal inputs due to alignment between their spectral bias and CFT smoothness.

  18. Grokking as Dimensional Phase Transition in Neural Networks

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Grokking occurs as the effective dimensionality of the gradient field transitions from sub-diffusive to super-diffusive at the onset of generalization, exhibiting self-organized criticality.

  19. Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues

    cond-mat.dis-nn 2026-02 conditional novelty 6.0 of 10

    A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.

  20. Quantitative Understanding of PDF Fits and their Uncertainties

    hep-ph 2025-12 conditional novelty 6.0 of 10

    After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.

  21. Sharpness-Aware Minimization for Efficiently Improving Generalization

    cs.LG 2020-10 conditional novelty 6.0 of 10

    SAM solves a min-max problem to locate flat low-loss regions, improving generalization on CIFAR, ImageNet and label-noise tasks.

  22. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  23. Theory of learning of high-dimensional controlled non-linear dynamical systems (I): models and methods

    cond-mat.dis-nn 2026-06 unverdicted novelty 5.0 of 10

    Introduces models for neural ODEs trained with online SGD and derives their high-dimensional learning curves via dynamical mean field theory.

  24. Bayesian Inference with Shaped Deep Non-linear MLPs

    math.ST 2026-05 unverdicted novelty 5.0 of 10

    In the LP/N = Θ(1) regime, Bayesian predictive posteriors for deep MLPs equal those of data-dependent kernels to first order, with a criterion identifying data processes that benefit from larger effective depth.

  25. Outer-Momentum Restarting in High-Dimensional Two-Phase Optimization

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Periodic outer-momentum restarts in two-phase optimizers exploit phase cancellation in a linearized NTK model to widen stable learning-rate and momentum ranges in language-model pretraining.

  26. The Thermodynamic Costs of Simple Linear Regression

    cond-mat.stat-mech 2026-05 unverdicted novelty 5.0 of 10

    Thermodynamic lower bounds are approximated for exact and SGD linear regression, producing energy-aware scaling laws for optimal training dataset size given a target generalization error.

  27. A physics-informed neural network approach to the point defect model for electrochemical oxide film growth

    cond-mat.mtrl-sci 2025-10 conditional novelty 5.0 of 10

    A hybrid PINN anchored by one FEM data point reproduces point-defect-model film thicknesses to about 1% error, while the pure PINN overpredicts by 2,400-5,700%.

  28. Is data-efficient learning feasible with quantum models?

    quant-ph 2025-08 conditional novelty 5.0 of 10

    Quantum kernels can beat an untuned classical kernel on datasets whose labels the authors deliberately construct from the quantum kernel's own spectrum, showing data efficiency by construction.

  29. Approximating Simple ReLU Networks based on Spectral Decomposition of Fisher Information

    stat.ML 2025-05 unverdicted novelty 5.0 of 10

    For random 2-layer ReLU networks the dominant eigenspaces of the Fisher information matrix are spanned by spherical harmonics of degree ≤2 and capture 97.7% of the trace independently of parameter count.

  30. Integrating Out, Twice:The Open-System Case That Neural-Network Ensemble Theory Is Missing

    cs.LG 2026-06 unverdicted novelty 4.0 of 10

    Neural-network ensembles match closed Gaussian systems but lack the open-system non-Hermitian generator and continuous spectrum required by nuclear optical models, yielding a structural negative on applicability.

  31. Neural Networks, Dispersion Relations and the Thermal Bootstrap

    hep-th 2026-05 unverdicted novelty 4.0 of 10

    A neural-network approach with dispersion relations handles infinite OPE towers in thermal conformal correlators without positivity.

  32. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

  33. Convergence rates for gradient descent in the training of overparameterized artificial neural networks with piecewise affine activation

    cs.LG 2021-02 unverdicted novelty 4.0 of 10

    Batch gradient descent achieves linear convergence to zero MSE with high probability for sufficiently wide shallow NNs with non-affine piecewise affine activations and distinct inputs.

  34. Lectures on Semiclassical Methods for Composite Operators

    hep-th 2026-06 unverdicted novelty 3.0 of 10

    Lecture notes develop semiclassical methods to compute large-n scaling dimensions of composite operators in CFTs, recovering known results in free theory and deriving one-loop corrections at the Wilson-Fisher fixed point.

  35. Some Inverse Problems in Particle Physics

    hep-lat 2026-06 unverdicted novelty 2.0 of 10

    Lectures reviewing three established numerical methods for inverse problems in extracting PDFs and spectral functions from lattice QCD and experimental data.

  36. Deep learning applied to computational mechanics: A comprehensive review, state of the art, and the classics

    cs.LG 2022-12 unverdicted novelty 2.0 of 10

    A comprehensive review of deep learning techniques for computational mechanics, including LSTM for constitutive modeling, PINNs for PDE solving, optimizers, and kernel methods.

Pith tools