REVIEW 2 cited by
Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Motivated by the studies of neural networks (e.g.,the neural tangent kernel theory), we perform a study on the large-dimensional behavior of kernel ridge regression (KRR) where the sample size $n \asymp d^{\gamma}$ for some $\gamma > 0$. Given an RKHS $\mathcal{H}$ associated with an inner product kernel defined on the sphere $\mathbb{S}^{d}$, we suppose that the true function $f_{\rho}^{*} \in [\mathcal{H}]^{s}$, the interpolation space of $\mathcal{H}$ with source condition $s>0$. We first determined the exact order (both upper and lower bound) of the generalization error of kernel ridge regression for the optimally chosen regularization parameter $\lambda$. We then further showed that when $0<s\le1$, KRR is minimax optimal; and when $s>1$, KRR is not minimax optimal (a.k.a. he saturation effect). Our results illustrate that the curves of rate varying along $\gamma$ exhibit the periodic plateau behavior and the multiple descent behavior and show how the curves evolve with $s>0$. Interestingly, our work provides a unified viewpoint of several recent works on kernel regression in the large-dimensional setting, which correspond to $s=0$ and $s=1$ respectively.
Forward citations
Cited by 2 Pith papers
-
Learning Curves of Stochastic Gradient Descent in Kernel Regression
Single-pass SGD with exponentially decaying steps is claimed to reach minimax-optimal excess risk in high-dimensional kernel regression for well-specified problems, with averaging handling misspecified problems.
-
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
For power-law data and target decays, SGD on the quadratically parameterized model provably beats linear SGD when the target opposes the spectrum, with rates T^{-(2β-2)/(α+β)} versus T^{-(β-1)/α}.
Discussion (0). Continue with ORCID to comment.