The Power of Randomization: Distributed Submodular Maximization on   Massive Datasets

Rafael da Ponte Barbosa; Alina Ene; Huy L. Nguyen; Justin; Ward

arXiv:1502.02606·cs.LG·April 23, 2015

The Power of Randomization: Distributed Submodular Maximization on Massive Datasets

Rafael da Ponte Barbosa, Alina Ene, Huy L. Nguyen, Justin, Ward

PDF

Open Access

TL;DR

This paper introduces a simple distributed algorithm for large-scale submodular maximization problems in machine learning, providing provable guarantees and near-centralized solution quality.

Contribution

It presents a distributed, embarrassingly parallel algorithm with theoretical approximation guarantees for constrained submodular maximization on massive datasets.

Findings

01

Algorithm achieves constant factor approximation guarantees.

02

Experimental results show near-centralized solution quality.

03

Efficiently handles large datasets with various constraints.

Abstract

A wide variety of problems in machine learning, including exemplar clustering, document summarization, and sensor placement, can be cast as constrained submodular maximization problems. Unfortunately, the resulting submodular optimization problems are often too large to be solved on a single machine. We develop a simple distributed algorithm that is embarrassingly parallel and it achieves provable, constant factor, worst-case approximation guarantees. In our experiments, we demonstrate its efficiency in large problems with different kinds of constraints with objective values always close to what is achievable in the centralized setting.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsPrivacy-Preserving Technologies in Data · Stochastic Gradient Optimization Techniques · Complexity and Algorithms in Graphs