A bi-criteria approximation algorithm for Means
Abstract: We consider the classical -means clustering problem in the setting bi-criteria approximation, in which an algoithm is allowed to output $\beta k > k$ clusters, and must produce a clustering with cost at most times the to the cost of the optimal set of clusters. We argue that this approach is natural in many settings, for which the exact number of clusters is a priori unknown, or unimportant up to a constant factor. We give new bi-criteria approximation algorithms, based on linear programming and local search, respectively, which attain a guarantee depending on the number of clusters that may be opened. Our gurantee is always at most and improves rapidly with (for example: $\alpha(2)<2.59$, and $\alpha(3) < 1.4$). Moreover, our algorithms have only polynomial dependence on the dimension of the input data, and so are applicable in high-dimensional settings.
Paper Prompts
Sign up for free to create and run prompts on this paper.