REVIEW 3 cited by
Row Sampling for Matrix Algorithms via a Non-Commutative Bernstein Bound
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We focus the use of \emph{row sampling} for approximating matrix algorithms. We give applications to matrix multipication; sparse matrix reconstruction; and, \math{\ell_2} regression. For a matrix \math{\matA\in\R^{m\times d}} which represents \math{m} points in \math{d\ll m} dimensions, all of these tasks can be achieved in \math{O(md^2)} via the singular value decomposition (SVD). For appropriate row-sampling probabilities (which typically depend on the norms of the rows of the \math{m\times d} left singular matrix of \math{\matA} (the \emph{leverage scores}), we give row-sampling algorithms with linear (up to polylog factors) dependence on the stable rank of \math{\matA}. This result is achieved through the application of non-commutative Bernstein bounds. We then give, to our knowledge, the first algorithms for computing approximations to the appropriate row-sampling probabilities without going through the SVD of \math{\matA}. Thus, these are the first \math{o(md^2)} algorithms for row-sampling based approximations to the matrix algorithms which use leverage scores as the sampling probabilities. The techniques we use to approximate sampling according to the leverage scores uses some powerful recent results in the theory of random projections for embedding, and may be of some independent interest. We confess that one may perform all these matrix tasks more efficiently using these same random projection methods, however the resulting algorithms are in terms of a small number of linear combinations of all the rows. In many applications, the actual rows of \math{\matA} have some physical meaning and so methods based on a small number of the actual rows are of interest.
Forward citations
Cited by 3 Pith papers
-
Fast, Space-Optimal Streaming Algorithms for Clustering and Subspace Embeddings
Streaming algorithms for (k,z)-clustering and Lp subspace embeddings can match offline algorithms in space and time, removing all dependence on stream length n.
-
Approximating the Top Eigenvector in Random Order Streams
A random-order streaming algorithm approximates the top eigenvector with near-linear memory whenever the spectral gap is constant, and a lower bound shows the heavy-row parameter is unavoidable.
-
On Socially Fair Low-Rank Approximation and Column Subset Selection
Fair low-rank approximation is NP-hard to approximate and requires exponential time under ETH; bicriteria algorithms achieve polynomial time with a rank and column count blow-up.
Discussion (0). Continue with ORCID to comment.