Pith. sign in

REVIEW

How Many Machines Can We Use in Parallel Computing for Kernel Ridge Regression?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.09948 v3 pith:AYMWLUI4 submitted 2018-05-25 math.ST stat.MLstat.TH

classification math.STstat.MLstat.TH
keywords regressionmachinescomputingestimationimportantkernelmanynonparametric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper aims to solve a basic problem in distributed statistical inference: how many machines can we use in parallel computing? In kernel ridge regression, we address this question in two important settings: nonparametric estimation and hypothesis testing. Specifically, we find a range for the number of machines under which optimal estimation/testing is achievable. The employed empirical processes method provides a unified framework, that allows us to handle various regression problems (such as thin-plate splines and nonparametric additive regression) under different settings (such as univariate, multivariate and diverging-dimensional designs). It is worth noting that the upper bounds of the number of machines are proven to be un-improvable (upto a logarithmic factor) in two important cases: smoothing spline regression and Gaussian RKHS regression. Our theoretical findings are backed by thorough numerical studies.

Discussion (0). Sign in to comment.

Pith tools