Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

Alexander Rakhlin; Tengyuan Liang

arxiv: 1808.00387 · v2 · pith:IEP3GLFWnew · submitted 2018-08-01 · 🧮 math.ST · cs.LG· stat.ML· stat.TH

Just Interpolate: Kernel "Ridgeless" Regression Can Generalize

Tengyuan Liang , Alexander Rakhlin This is my paper

classification 🧮 math.ST cs.LGstat.MLstat.TH

keywords datakernelgeneralizeinterpolatedphenomenonregressionregularizationridgeless

0 comments

read the original abstract

In the absence of explicit regularization, Kernel "Ridgeless" Regression with nonlinear kernels has the potential to fit the training data perfectly. It has been observed empirically, however, that such interpolated solutions can still generalize well on test data. We isolate a phenomenon of implicit regularization for minimum-norm interpolated solutions which is due to a combination of high dimensionality of the input data, curvature of the kernel function, and favorable geometric properties of the data such as an eigenvalue decay of the empirical covariance and kernel matrices. In addition to deriving a data-dependent upper bound on the out-of-sample error, we present experimental evidence suggesting that the phenomenon occurs in the MNIST dataset.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

When Does $\ell_2$-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the $\ell_1$ Implicit Bias
cs.LG 2026-05 unverdicted novelty 8.0

ℓ₂-Boosting exhibits benign overfitting with logarithmic excess variance decay Θ(σ²/log(p/n)) under isotropic noise due to ℓ₁ bias, and a subdifferential early stopping rule recovers minimax-optimal ℓ₁ rates.
When Does $\ell_2$-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the $\ell_1$ Implicit Bias
cs.LG 2026-05 unverdicted novelty 7.0

ℓ₂-boosting localizes noise into sparse sets under isotropic pure-noise models, yielding excess variance Θ(σ²/log(p/n)) instead of linear decay, with a tuning-free early stopping rule attaining minimax ℓ₁ rates.
A Ridge Too Far: Correcting Over-Shrinkage via Negative Regularization
cs.LG 2025-08 unverdicted novelty 6.0

Negative-capable ridge regression uses controlled negative regularization as anti-shrinkage to increase effective complexity along weak eigendirections and mitigate underfitting in small-data regression.