Pith. sign in

REVIEW 2 cited by

Optimal Rate of Kernel Regression in Large Dimensions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.04268 v2 pith:NI6N47WU submitted 2023-09-08 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH
keywords gammakernelrateregressionoptimalasympbehaviorbound
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We perform a study on kernel regression for large-dimensional data (where the sample size $n$ is polynomially depending on the dimension $d$ of the samples, i.e., $n\asymp d^{\gamma}$ for some $\gamma >0$ ). We first build a general tool to characterize the upper bound and the minimax lower bound of kernel regression for large dimensional data through the Mendelson complexity $\varepsilon_{n}^{2}$ and the metric entropy $\bar{\varepsilon}_{n}^{2}$ respectively. When the target function falls into the RKHS associated with a (general) inner product model defined on $\mathbb{S}^{d}$, we utilize the new tool to show that the minimax rate of the excess risk of kernel regression is $n^{-1/2}$ when $n\asymp d^{\gamma}$ for $\gamma =2, 4, 6, 8, \cdots$. We then further determine the optimal rate of the excess risk of kernel regression for all the $\gamma>0$ and find that the curve of optimal rate varying along $\gamma$ exhibits several new phenomena including the multiple descent behavior and the periodic plateau behavior. As an application, For the neural tangent kernel (NTK), we also provide a similar explicit description of the curve of optimal rate. As a direct corollary, we know these claims hold for wide neural networks as well.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Curves of Stochastic Gradient Descent in Kernel Regression

    stat.ML 2025-05 reject novelty 7.0 of 10

    Single-pass SGD with exponentially decaying steps is claimed to reach minimax-optimal excess risk in high-dimensional kernel regression for well-specified problems, with averaging handling misspecified problems.

  2. Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression

    cs.LG 2025-02 conditional novelty 7.0 of 10

    For power-law data and target decays, SGD on the quadratically parameterized model provably beats linear SGD when the target opposes the spectrum, with rates T^{-(2β-2)/(α+β)} versus T^{-(β-1)/α}.

Pith tools