Pith. sign in

REVIEW 2 cited by

Distributed Estimation of Principal Eigenspaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1702.06488 v4 pith:YSYFI5CY submitted 2017-02-21 stat.CO math.STstat.TH

classification stat.COmath.STstat.TH
keywords distributedmachinescentraldataprincipalacrossanalysiscovariance
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Principal component analysis (PCA) is fundamental to statistical machine learning. It extracts latent principal factors that contribute to the most variation of the data. When data are stored across multiple machines, however, communication cost can prohibit the computation of PCA in a central location and distributed algorithms for PCA are thus needed. This paper proposes and studies a distributed PCA algorithm: each node machine computes the top $K$ eigenvectors and transmits them to the central server; the central server then aggregates the information from all the node machines and conducts a PCA based on the aggregated information. We investigate the bias and variance for the resulting distributed estimator of the top $K$ eigenvectors. In particular, we show that for distributions with symmetric innovation, the empirical top eigenspaces are unbiased and hence the distributed PCA is "unbiased". We derive the rate of convergence for distributed PCA estimators, which depends explicitly on the effective rank of covariance, eigen-gap, and the number of machines. We show that when the number of machines is not unreasonably large, the distributed PCA performs as well as the whole sample PCA, even without full access of whole data. The theoretical results are verified by an extensive simulation study. We also extend our analysis to the heterogeneous case where the population covariance matrices are different across local machines but share similar top eigen-structures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rates of Convergence for Large-scale Nearest Neighbor Classification

    stat.ML 2019-09 conditional novelty 6.0 of 10

    A distributed k-nearest-neighbor classifier that pools local predictions by majority vote attains the same minimax-optimal excess risk and instability rates as the oracle full-data kNN classifier.

  2. Least Squares Approximation for a Distributed System

    stat.ME 2019-08 conditional novelty 5.0 of 10

    A distributed least squares approximation combines local estimators weighted by inverse covariance to match global estimator efficiency with one communication round.

Pith tools