Papers
Topics
Authors
Recent
Search
2000 character limit reached

Fast and Sample Near-Optimal Algorithms for Learning Multidimensional Histograms

Published 23 Feb 2018 in cs.LG, cs.DS, math.ST, and stat.TH | (1802.08513v1)

Abstract: We study the problem of robustly learning multi-dimensional histograms. A dd-dimensional function h:D→Rh: D \rightarrow \mathbb{R} is called a kk-histogram if there exists a partition of the domain D⊆R<sup>dD \subseteq \mathbb{R}<sup>d into kk axis-aligned rectangles such that hh is constant within each such rectangle. Let f:D→Rf: D \rightarrow \mathbb{R} be a dd-dimensional probability density function and suppose that ff is OPT\mathrm{OPT}-close, in L1L_1-distance, to an unknown kk-histogram (with unknown partition). Our goal is to output a hypothesis that is O(OPT)+ϵO(\mathrm{OPT}) + \epsilon close to ff, in L1L_1-distance. We give an algorithm for this learning problem that uses n=O~d(k/ϵ<sup>2)n = \tilde{O}_d(k/\epsilon<sup>2) samples and runs in time O~d(n)\tilde{O}_d(n). For any fixed dimension, our algorithm has optimal sample complexity, up to logarithmic factors, and runs in near-linear time. Prior to our work, the time complexity of the d=1d=1 case was well-understood, but significant gaps in our understanding remained even for d=2d=2.

Citations (19)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.