The Power of Randomization: Distributed Submodular Maximization on Massive Datasets

Alina Ene; Huy L. Nguyen; Justin Ward; Rafael da Ponte Barbosa

arxiv: 1502.02606 · v2 · pith:TZY3WCICnew · submitted 2015-02-09 · 💻 cs.LG · cs.AI· cs.DC

The Power of Randomization: Distributed Submodular Maximization on Massive Datasets

Rafael da Ponte Barbosa , Alina Ene , Huy L. Nguyen , Justin Ward This is my paper

classification 💻 cs.LG cs.AIcs.DC

keywords problemssubmodulardistributedlargemachinemaximizationachievableachieves

0 comments

read the original abstract

A wide variety of problems in machine learning, including exemplar clustering, document summarization, and sensor placement, can be cast as constrained submodular maximization problems. Unfortunately, the resulting submodular optimization problems are often too large to be solved on a single machine. We develop a simple distributed algorithm that is embarrassingly parallel and it achieves provable, constant factor, worst-case approximation guarantees. In our experiments, we demonstrate its efficiency in large problems with different kinds of constraints with objective values always close to what is achievable in the centralized setting.

This paper has not been read by Pith yet.

The Power of Randomization: Distributed Submodular Maximization on Massive Datasets

discussion (0)