Comparative Analysis of Optimization Strategies for K-means Clustering   in Big Data Contexts: A Review

Ravil Mussabayev; Rustam Mussabayev

arXiv:2310.09819·cs.LG·May 21, 2024·1 cites

Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review

Ravil Mussabayev, Rustam Mussabayev

PDF

Open Access

TL;DR

This paper reviews various optimization strategies for K-means clustering in big data, comparing their performance on large datasets to guide practitioners in selecting suitable methods based on speed, quality, and simplicity.

Contribution

It provides a comprehensive comparison of optimization techniques for K-means in big data, highlighting trade-offs and practical insights for improved scalability.

Findings

01

Different techniques excel on different dataset types.

02

Trade-offs exist between speed and clustering accuracy.

03

Sampling and approximation methods improve scalability.

Abstract

This paper presents a comparative analysis of different optimization techniques for the K-means algorithm in the context of big data. K-means is a widely used clustering algorithm, but it can suffer from scalability issues when dealing with large datasets. The paper explores different approaches to overcome these issues, including parallelization, approximation, and sampling methods. The authors evaluate the performance of various clustering techniques on a large number of benchmark datasets, comparing them according to the dominance criterion provided by the "less is more" approach (LIMA), i.e., simultaneously along the dimensions of speed, clustering quality, and simplicity. The results show that different techniques are more suitable for different types of datasets and provide insights into the trade-offs between speed and accuracy in K-means clustering for big data. Overall, the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsData Stream Mining Techniques · Advanced Clustering Algorithms Research · Metaheuristic Optimization Algorithms Research

Methodsk-Means Clustering · SPEED: Separable Pyramidal Pooling EncodEr-Decoder for Real-Time Monocular Depth Estimation on Low-Resource Settings