The paper reports numerical evidence that the NTK equivalence theorem does not hold for finite-width neural networks trained with SGD, because trained networks and NTK kernel regressors behave differently when a layer is added.
Kernel methods for deep learning // Advances in neural information processing systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Issues with Neural Tangent Kernel Approach to Neural Networks
The paper reports numerical evidence that the NTK equivalence theorem does not hold for finite-width neural networks trained with SGD, because trained networks and NTK kernel regressors behave differently when a layer is added.