Pith. sign in

REVIEW 2 cited by

Optimal Convergence for Distributed Learning with Stochastic Gradient Methods and Spectral Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.07226 v2 pith:XKIX3QN4 submitted 2018-01-22 stat.ML cs.AIcs.LGmath.FA

classification stat.MLcs.AIcs.LGmath.FA
keywords distributedalgorithmsgradientkernelmethodsoptimalregressionresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study generalization properties of distributed algorithms in the setting of nonparametric regression over a reproducing kernel Hilbert space (RKHS). We first investigate distributed stochastic gradient methods (SGM), with mini-batches and multi-passes over the data. We show that optimal generalization error bounds can be retained for distributed SGM provided that the partition level is not too large. We then extend our results to spectral-regularization algorithms (SRA), including kernel ridge regression (KRR), kernel principal component analysis, and gradient methods. Our results are superior to the state-of-the-art theory. Particularly, our results show that distributed SGM has a smaller theoretical computational complexity, compared with distributed KRR and classic SGM. Moreover, even for non-distributed SRA, they provide the first optimal, capacity-dependent convergence rates, considering the case that the regression function may not be in the RKHS.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Random feature approximation for general spectral methods

    stat.ML 2025-06 conditional novelty 6.0 of 10

    Under source conditions with smoothness r>0 and capacity 2r+b>1, random features achieve minimax-optimal rates for any spectral regularization method with qualification at least r∨1.

  2. Optimal Convergence Rates for Neural Operators

    stat.ML 2024-12 conditional novelty 5.0 of 10

    Two-layer neural operators trained with early-stopped gradient descent achieve the same minimax convergence rates as kernel methods in the neural tangent kernel regime.

Pith tools