Papers
Topics
Authors
Recent
Search
2000 character limit reached

Kernel Thinning

Published 12 May 2021 in stat.ML, cs.LG, math.ST, stat.CO, stat.ME, and stat.TH | (2105.05842v11)

Abstract: We introduce kernel thinning, a new procedure for compressing a distribution P\mathbb{P} more effectively than i.i.d. sampling or standard thinning. Given a suitable reproducing kernel k\mathbf{k}_{\star} and O(n<sup>2)O(n<sup>2) time, kernel thinning compresses an nn-point approximation to P\mathbb{P} into a n\sqrt{n}-point approximation with comparable worst-case integration error across the associated reproducing kernel Hilbert space. The maximum discrepancy in integration error is Od(n<sup>1/2log</sup>n)O_d(n<sup>{-1/2}\sqrt{\log</sup> n}) in probability for compactly supported P\mathbb{P} and Od(n<sup>12</sup>(logn)<sup>(d+1)/2loglog</sup>n)O_d(n<sup>{-\frac{1}{2}}</sup> (\log n)<sup>{(d+1)/2}\sqrt{\log\log</sup> n}) for sub-exponential P\mathbb{P} on R<sup>d\mathbb{R}<sup>d. In contrast, an equal-sized i.i.d. sample from P\mathbb{P} suffers Ω(n<sup>1/4)\Omega(n<sup>{-1/4}) integration error. Our sub-exponential guarantees resemble the classical quasi-Monte Carlo error rates for uniform P\mathbb{P} on [0,1]<sup>d[0,1]<sup>d but apply to general distributions on R<sup>d\mathbb{R}<sup>d and a wide range of common kernels. Moreover, the same construction delivers near-optimal L<sup>L<sup>\infty coresets in O(n<sup>2)O(n<sup>2) time. We use our results to derive explicit non-asymptotic maximum mean discrepancy bounds for Gaussian, Mat\'ern, and B-spline kernels and present two vignettes illustrating the practical benefits of kernel thinning over i.i.d. sampling and standard Markov chain Monte Carlo thinning, in dimensions d=2d=2 through $100$.

Authors (2)
Citations (32)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.