Distributed Estimation for Principal Component Analysis: an Enlarged   Eigenspace Analysis

Xi Chen; Jason D. Lee; He Li; Yun Yang

arXiv:2004.02336·cs.DC·February 4, 2021·5 cites

Distributed Estimation for Principal Component Analysis: an Enlarged Eigenspace Analysis

Xi Chen, Jason D. Lee, He Li, Yun Yang

PDF

Open Access

TL;DR

This paper introduces a communication-efficient distributed algorithm for estimating the top-L eigenspace in PCA without requiring eigengap assumptions, suitable for large-scale data analysis.

Contribution

A novel multi-round distributed PCA algorithm that leverages shift-and-invert preconditioning and convex optimization, removing restrictions on the number of machines.

Findings

01

Achieves fast convergence rate in distributed setting

02

No restriction on the number of machines for the algorithm

03

Demonstrates effectiveness through simulation studies

Abstract

The growing size of modern data sets brings many challenges to the existing statistical estimation approaches, which calls for new distributed methodologies. This paper studies distributed estimation for a fundamental statistical machine learning problem, principal component analysis (PCA). Despite the massive literature on top eigenvector estimation, much less is presented for the top- $L$ -dim ( $L > 1$ ) eigenspace estimation, especially in a distributed manner. We propose a novel multi-round algorithm for constructing top- $L$ -dim eigenspace for distributed data. Our algorithm takes advantage of shift-and-invert preconditioning and convex optimization. Our estimator is communication-efficient and achieves a fast convergence rate. In contrast to the existing divide-and-conquer algorithm, our approach has no restriction on the number of machines. Theoretically, the traditional Davis-Kahan…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSparse and Compressive Sensing Techniques · Random Matrices and Applications · Statistical Methods and Inference

MethodsPrincipal Components Analysis