Pith. sign in

REVIEW 3 cited by

To understand deep learning we need to understand kernel learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.01396 v3 pith:IOY6NWHY submitted 2018-02-05 stat.ML cs.LG

classification stat.MLcs.LG
keywords deepkernellearningdataclassifiersgeneralizationkernelsmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generalization performance of classifiers in deep learning has recently become a subject of intense study. Deep models, typically over-parametrized, tend to fit the training data exactly. Despite this "overfitting", they perform well on test data, a phenomenon not yet fully understood. The first point of our paper is that strong performance of overfitted classifiers is not a unique feature of deep learning. Using six real-world and two synthetic datasets, we establish experimentally that kernel machines trained to have zero classification or near zero regression error perform very well on test data, even when the labels are corrupted with a high level of noise. We proceed to give a lower bound on the norm of zero loss solutions for smooth kernels, showing that they increase nearly exponentially with data size. We point out that this is difficult to reconcile with the existing generalization bounds. Moreover, none of the bounds produce non-trivial results for interpolating solutions. Second, we show experimentally that (non-smooth) Laplacian kernels easily fit random labels, a finding that parallels results for ReLU neural networks. In contrast, fitting noisy data requires many more epochs for smooth Gaussian kernels. Similar performance of overfitted Laplacian and Gaussian classifiers on test, suggests that generalization is tied to the properties of the kernel function rather than the optimization process. Certain key phenomena of deep learning are manifested similarly in kernel methods in the modern "overfitted" regime. The combination of the experimental and theoretical results presented in this paper indicates a need for new theoretical ideas for understanding properties of classical kernel methods. We argue that progress on understanding deep learning will be difficult until more tractable "shallow" kernel methods are better understood.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 110 citations worldwide. Full citation record

  1. On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels

    math.ST 2019-08 conditional novelty 8.0 of 10

    Minimum-norm interpolants in reproducing kernel Hilbert spaces have risk that can exhibit multiple peaks and valleys as the sample size grows, with peak locations predicted by the scaling d = n^α.

  2. The generalization error of random features regression: Precise asymptotics and double descent curve

    math.ST 2019-08 conditional novelty 8.0 of 10

    Mei and Montanari derive the exact asymptotic test error of random features ridge regression and show it reproduces the full double descent phenomenon without any misspecified structure.

  3. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 2.0 of 10

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.

Pith tools