Papers
Topics
Authors
Recent
Search
2000 character limit reached

A simple D^2-sampling based PTAS for k-means and other Clustering Problems

Published 20 Jan 2012 in cs.DS | (1201.4206v1)

Abstract: Given a set of points P⊂R<sup>dP \subset \mathbb{R}<sup>d, the kk-means clustering problem is to find a set of kk {\em centers} C=c1,...,ck,ci∈R<sup>d,C = {c_1,...,c_k}, c_i \in \mathbb{R}<sup>d, such that the objective function ∑x∈Pd(x,C)<sup>2\sum_{x \in P} d(x,C)<sup>2, where d(x,C)d(x,C) denotes the distance between xx and the closest center in CC, is minimized. This is one of the most prominent objective functions that have been studied with respect to clustering. D<sup>2D<sup>2-sampling \cite{ArthurV07} is a simple non-uniform sampling technique for choosing points from a set of points. It works as follows: given a set of points P⊆R<sup>dP \subseteq \mathbb{R}<sup>d, the first point is chosen uniformly at random from PP. Subsequently, a point from PP is chosen as the next sample with probability proportional to the square of the distance of this point to the nearest previously sampled points. D<sup>2D<sup>2-sampling has been shown to have nice properties with respect to the kk-means clustering problem. Arthur and Vassilvitskii \cite{ArthurV07} show that kk points chosen as centers from PP using D<sup>2D<sup>2-sampling gives an O(log⁡k)O(\log{k}) approximation in expectation. Ailon et. al. \cite{AJMonteleoni09} and Aggarwal et. al. \cite{AggarwalDK09} extended results of \cite{ArthurV07} to show that O(k)O(k) points chosen as centers using D<sup>2D<sup>2-sampling give O(1)O(1) approximation to the kk-means objective function with high probability. In this paper, we further demonstrate the power of D<sup>2D<sup>2-sampling by giving a simple randomized (1+ϵ)(1 + \epsilon)-approximation algorithm that uses the D<sup>2D<sup>2-sampling in its core.

Citations (61)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.