REVIEW 2 cited by
Optimal Rate of Kernel Regression in Large Dimensions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We perform a study on kernel regression for large-dimensional data (where the sample size $n$ is polynomially depending on the dimension $d$ of the samples, i.e., $n\asymp d^{\gamma}$ for some $\gamma >0$ ). We first build a general tool to characterize the upper bound and the minimax lower bound of kernel regression for large dimensional data through the Mendelson complexity $\varepsilon_{n}^{2}$ and the metric entropy $\bar{\varepsilon}_{n}^{2}$ respectively. When the target function falls into the RKHS associated with a (general) inner product model defined on $\mathbb{S}^{d}$, we utilize the new tool to show that the minimax rate of the excess risk of kernel regression is $n^{-1/2}$ when $n\asymp d^{\gamma}$ for $\gamma =2, 4, 6, 8, \cdots$. We then further determine the optimal rate of the excess risk of kernel regression for all the $\gamma>0$ and find that the curve of optimal rate varying along $\gamma$ exhibits several new phenomena including the multiple descent behavior and the periodic plateau behavior. As an application, For the neural tangent kernel (NTK), we also provide a similar explicit description of the curve of optimal rate. As a direct corollary, we know these claims hold for wide neural networks as well.
Forward citations
Cited by 2 Pith papers
-
Learning Curves of Stochastic Gradient Descent in Kernel Regression
Single-pass SGD with exponentially decaying steps is claimed to reach minimax-optimal excess risk in high-dimensional kernel regression for well-specified problems, with averaging handling misspecified problems.
-
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
For power-law data and target decays, SGD on the quadratically parameterized model provably beats linear SGD when the target opposes the spectrum, with rates T^{-(2β-2)/(α+β)} versus T^{-(β-1)/α}.
Discussion (0). Continue with ORCID to comment.