A tight lower bound instance for k-means++ in constant dimension (1401.2912v2)

Published 13 Jan 2014 in cs.DS

Abstract: The k-means++ seeding algorithm is one of the most popular algorithms that is used for finding the initial $k$ centers when using the k-means heuristic. The algorithm is a simple sampling procedure and can be described as follows: Pick the first center randomly from the given points. For $i > 1$, pick a point to be the $i^{th}$ center with probability proportional to the square of the Euclidean distance of this point to the closest previously $(i-1)$ chosen centers. The k-means++ seeding algorithm is not only simple and fast but also gives an $O(\log{k})$ approximation in expectation as shown by Arthur and Vassilvitskii. There are datasets on which this seeding algorithm gives an approximation factor of $\Omega(\log{k})$ in expectation. However, it is not clear from these results if the algorithm achieves good approximation factor with reasonably high probability (say $1/poly(k)$). Brunsch and R\"{o}glin gave a dataset where the k-means++ seeding algorithm achieves an $O(\log{k})$ approximation ratio with probability that is exponentially small in $k$. However, this and all other known lower-bound examples are high dimensional. So, an open problem was to understand the behavior of the algorithm on low dimensional datasets. In this work, we give a simple two dimensional dataset on which the seeding algorithm achieves an $O(\log{k})$ approximation ratio with probability exponentially small in $k$. This solves open problems posed by Mahajan et al. and by Brunsch and R\"{o}glin.

Authors (3)

Anup Bhattacharya (15 papers)
Ragesh Jaiswal (27 papers)
Nir Ailon (31 papers)

Citations (10)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Related Papers

A bad 2-dimensional instance for k-means++ (2013)
Improved Outlier Robust Seeding for k-means (2023)
Fast and Accurate $k$-means++ via Rejection Sampling (2020)
Noisy, Greedy and Not So Greedy k-means++ (2019)
Seeding K-Means using Method of Moments (2015)