REVIEW 2 cited by
Generalization in Kernel Regression Under Realistic Assumptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
It is by now well-established that modern over-parameterized models seem to elude the bias-variance tradeoff and generalize well despite overfitting noise. Many recent works attempt to analyze this phenomenon in the relatively tractable setting of kernel regression. However, as we argue in detail, most past works on this topic either make unrealistic assumptions, or focus on a narrow problem setup. This work aims to provide a unified theory to upper bound the excess risk of kernel regression for nearly all common and realistic settings. Specifically, we provide rigorous bounds that hold for common kernels and for any amount of regularization, noise, any input dimension, and any number of samples. Furthermore, we provide relative perturbation bounds for the eigenvalues of kernel matrices, which may be of independent interest. These reveal a self-regularization phenomenon, whereby a heavy tail in the eigendecomposition of the kernel provides it with an implicit form of regularization, enabling good generalization. When applied to common kernels, our results imply benign overfitting in high input dimensions, nearly tempered overfitting in fixed dimensions, and explicit convergence rates for regularized regression. As a by-product, we obtain time-dependent bounds for neural networks trained in the kernel regime.
Forward citations
Cited by 2 Pith papers
-
Interpretable QSPR Modeling using Recursive Feature Machines and Multi-scale Fingerprints
Using Recursive Feature Machines with a custom hybrid fingerprint yields lower solubility prediction errors than graph neural networks on ESOL and FreeSolv, while also producing feature-importance scores.
-
Towards a Statistical Understanding of Neural Networks: Beyond the Neural Tangent Kernel Theories
The paper reviews fixed-kernel neural network theory and proposes an over-parameterized Gaussian sequence model as a prototype for feature learning.
Discussion (0). Continue with ORCID to comment.