Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sets Clustering

Published 9 Mar 2020 in cs.LG and stat.ML | (2003.04135v1)

Abstract: The input to the \emph{sets-kk-means} problem is an integer k≥1k\geq 1 and a set P=P1,⋯ ,Pn\mathcal{P}={P_1,\cdots,P_n} of sets in R<sup>d\mathbb{R}<sup>d. The goal is to compute a set CC of kk centers (points) in R<sup>d\mathbb{R}<sup>d that minimizes the sum ∑P∈Pmin⁡p∈P,c∈C∣p−c∣<sup>2\sum_{P\in \mathcal{P}} \min_{p\in P, c\in C}\left| p-c \right|<sup>2 of squared distances to these sets. An \emph{ε\varepsilon-core-set} for this problem is a weighted subset of P\mathcal{P} that approximates this sum up to 1±ε1\pm\varepsilon factor, for \emph{every} set CC of kk centers in R<sup>d\mathbb{R}<sup>d. We prove that such a core-set of O(log⁡<sup>2n)O(\log<sup>2{n}) sets always exists, and can be computed in O(nlog⁡n)O(n\log{n}) time, for every input P\mathcal{P} and every fixed d,k≥1d,k\geq 1 and ε∈(0,1)\varepsilon \in (0,1). The result easily generalized for any metric space, distances to the power of $z&gt;0$, and M-estimators that handle outliers. Applying an inefficient but optimal algorithm on this coreset allows us to obtain the first PTAS (1+ε1+\varepsilon approximation) for the sets-kk-means problem that takes time near linear in nn. This is the first result even for sets-mean on the plane (k=1k=1, d=2d=2). Open source code and experimental results for document classification and facility locations are also provided.

Citations (21)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.