Pith. sign in

REVIEW 12 cited by

Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.07572 v4 pith:DPSOYAJ3 submitted 2018-06-20 cs.LG cs.NEmath.PRstat.ML

classification cs.LGcs.NEmath.PRstat.ML
keywords kerneltrainingduringinfinite-widthlimitneuralannsconvergence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

At initialization, artificial neural networks (ANNs) are equivalent to Gaussian processes in the infinite-width limit, thus connecting them to kernel methods. We prove that the evolution of an ANN during training can also be described by a kernel: during gradient descent on the parameters of an ANN, the network function $f_\theta$ (which maps input vectors to output vectors) follows the kernel gradient of the functional cost (which is convex, in contrast to the parameter cost) w.r.t. a new kernel: the Neural Tangent Kernel (NTK). This kernel is central to describe the generalization features of ANNs. While the NTK is random at initialization and varies during training, in the infinite-width limit it converges to an explicit limiting kernel and it stays constant during training. This makes it possible to study the training of ANNs in function space instead of parameter space. Convergence of the training can then be related to the positive-definiteness of the limiting NTK. We prove the positive-definiteness of the limiting NTK when the data is supported on the sphere and the non-linearity is non-polynomial. We then focus on the setting of least-squares regression and show that in the infinite-width limit, the network function $f_\theta$ follows a linear differential equation during training. The convergence is fastest along the largest kernel principal components of the input data with respect to the NTK, hence suggesting a theoretical motivation for early stopping. Finally we study the NTK numerically, observe its behavior for wide networks, and compare it to the infinite-width limit.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neural Spectral Bias and Conformal Correlators I: Introduction and Applications

    hep-th 2026-04 unverdicted novelty 8.0 of 10

    Simple feed-forward neural networks trained on crossing symmetry plus a single anchor value reproduce CFT correlators to percent-level accuracy, and the authors conjecture this works because physical correlators are t...

  2. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  3. Coupled by Design: Computing Kerr-Newman Quasinormal Modes with a Hybrid SpectralPINN Solver

    gr-qc 2026-07 conditional novelty 6.0 of 10

    A hybrid spectral/PINN solver produces the first systematic public dataset of Kerr-Newman quasinormal-mode frequencies across the full sub-extremal parameter space, including both gravitational- and vector-led branches.

  4. Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A task ma+nb mod p is representable by a z^k holomorphic network iff m+n=k; non-representable tasks cannot be memorised at any width.

  5. Neural Networks Reveal a Universal Bias in Conformal Correlators

    hep-th 2026-04 unverdicted novelty 6.0 of 10

    Simple neural networks trained on crossing symmetry and one anchor point reproduce conformal correlators to within a few percent across many CFTs.

  6. Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvalues

    cond-mat.dis-nn 2026-02 conditional novelty 6.0 of 10

    A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.

  7. Quantitative Understanding of PDF Fits and their Uncertainties

    hep-ph 2025-12 conditional novelty 6.0 of 10

    After an initial transient, a PDF-fitting neural network's output obeys f_t = U(t) f_0 + V(t) Y, a linear blend of the initial network and the data with explicit time-dependent operators.

  8. Pre-Strings Lectures on Artificial Intelligence

    hep-th 2026-07 accept novelty 5.5 of 10

    Lecture notes define neural-network field theory and survey how it recovers known QFT/string results plus applied AI techniques for string problems.

  9. A physics-informed neural network approach to the point defect model for electrochemical oxide film growth

    cond-mat.mtrl-sci 2025-10 conditional novelty 5.0 of 10

    A hybrid PINN anchored by one FEM data point reproduces point-defect-model film thicknesses to about 1% error, while the pure PINN overpredicts by 2,400-5,700%.

  10. Is data-efficient learning feasible with quantum models?

    quant-ph 2025-08 conditional novelty 5.0 of 10

    Quantum kernels can beat an untuned classical kernel on datasets whose labels the authors deliberately construct from the quantum kernel's own spectrum, showing data efficiency by construction.

  11. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

  12. Statistical Machine Learning for Astronomy -- A Textbook

    astro-ph.IM 2025-06 unverdicted

    A systematic Bayesian-first textbook that derives classical and modern machine learning methods for astronomy from probability theory, explicitly without new research results.

Pith tools