Pith. sign in

REVIEW 2 cited by

Optimal Rates of Kernel Ridge Regression under Source Condition in Large Dimensions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01270 v1 pith:GAMUQXPA submitted 2024-01-02 cs.LG

classification cs.LG
keywords kernelregressionbehaviorgammamathcaloptimalridgecondition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Motivated by the studies of neural networks (e.g.,the neural tangent kernel theory), we perform a study on the large-dimensional behavior of kernel ridge regression (KRR) where the sample size $n \asymp d^{\gamma}$ for some $\gamma > 0$. Given an RKHS $\mathcal{H}$ associated with an inner product kernel defined on the sphere $\mathbb{S}^{d}$, we suppose that the true function $f_{\rho}^{*} \in [\mathcal{H}]^{s}$, the interpolation space of $\mathcal{H}$ with source condition $s>0$. We first determined the exact order (both upper and lower bound) of the generalization error of kernel ridge regression for the optimally chosen regularization parameter $\lambda$. We then further showed that when $0<s\le1$, KRR is minimax optimal; and when $s>1$, KRR is not minimax optimal (a.k.a. he saturation effect). Our results illustrate that the curves of rate varying along $\gamma$ exhibit the periodic plateau behavior and the multiple descent behavior and show how the curves evolve with $s>0$. Interestingly, our work provides a unified viewpoint of several recent works on kernel regression in the large-dimensional setting, which correspond to $s=0$ and $s=1$ respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Curves of Stochastic Gradient Descent in Kernel Regression

    stat.ML 2025-05 reject novelty 7.0 of 10

    Single-pass SGD with exponentially decaying steps is claimed to reach minimax-optimal excess risk in high-dimensional kernel regression for well-specified problems, with averaging handling misspecified problems.

  2. Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression

    cs.LG 2025-02 conditional novelty 7.0 of 10

    For power-law data and target decays, SGD on the quadratically parameterized model provably beats linear SGD when the target opposes the spectrum, with rates T^{-(2β-2)/(α+β)} versus T^{-(β-1)/α}.

Pith tools