An Empirical Evaluation of $k$-Means Coresets

Chris Schwiegelshohn; Omar Ali Sheikh-Omar

arXiv:2207.00966·cs.DS·July 5, 2022

An Empirical Evaluation of $k$-Means Coresets

Chris Schwiegelshohn, Omar Ali Sheikh-Omar

PDF

1 Repo

TL;DR

This paper evaluates the quality of various $k$-means coresets using a new benchmark and real-world data, revealing the challenges in measuring coreset distortion and providing practical insights.

Contribution

It introduces a benchmark for evaluating $k$-means coresets and performs an extensive empirical comparison of existing algorithms.

Findings

01

Coreset quality varies significantly across algorithms.

02

Measuring coreset distortion is computationally challenging.

03

The benchmark facilitates heuristic evaluation of coresets.

Abstract

Coresets are among the most popular paradigms for summarizing data. In particular, there exist many high performance coresets for clustering problems such as $k$ -means in both theory and practice. Curiously, there exists no work on comparing the quality of available $k$ -means coresets. In this paper we perform such an evaluation. There currently is no algorithm known to measure the distortion of a candidate coreset. We provide some evidence as to why this might be computationally difficult. To complement this, we propose a benchmark for which we argue that computing coresets is challenging and which also allows us an easy (heuristic) evaluation of coresets. Using this benchmark and real-world data sets, we conduct an exhaustive evaluation of the most commonly used coreset algorithms from theory and practice.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

sheikhomar/eval-k-means-coresets
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsCoresets