Pith. sign in

REVIEW 3 cited by

Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.00387 v2 pith:IEP3GLFW submitted 2018-08-01 math.ST cs.LGstat.MLstat.TH

classification math.STcs.LGstat.MLstat.TH
keywords datakernelgeneralizeinterpolatedphenomenonregressionregularizationridgeless
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In the absence of explicit regularization, Kernel "Ridgeless" Regression with nonlinear kernels has the potential to fit the training data perfectly. It has been observed empirically, however, that such interpolated solutions can still generalize well on test data. We isolate a phenomenon of implicit regularization for minimum-norm interpolated solutions which is due to a combination of high dimensionality of the input data, curvature of the kernel function, and favorable geometric properties of the data such as an eigenvalue decay of the empirical covariance and kernel matrices. In addition to deriving a data-dependent upper bound on the out-of-sample error, we present experimental evidence suggesting that the phenomenon occurs in the MNIST dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The generalization error of random features regression: Precise asymptotics and double descent curve

    math.ST 2019-08 conditional novelty 8.0 of 10

    Mei and Montanari derive the exact asymptotic test error of random features ridge regression and show it reproduces the full double descent phenomenon without any misspecified structure.

  2. The Cost of Discretization in Functional Linear Regression: Minimax Rates and Adaptation

    math.ST 2026-07 accept novelty 7.0 of 10

    Matching minimax prediction rates for discretely observed functional linear regression are n^{-ν/(ν+1)}+(nm)^{-ν/κ} under independent design, and those two terms plus m^{-ν}+m^{-4α} under common design.

  3. Deep neural networks, generic universal interpolation, and controlled ODEs

    math.OC 2019-08 accept novelty 7.0 of 10

    Finite training sets of any size can be exactly interpolated by a controlled ODE with five fixed vector fields, and this property holds generically for random real analytic vector fields.

Pith tools