Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning kk-Modal Distributions via Testing

Published 13 Jul 2011 in cs.DS, cs.LG, math.ST, and stat.TH | (1107.2700v3)

Abstract: A kk-modal probability distribution over the discrete domain 1,...,n{1,...,n} is one whose histogram has at most kk "peaks" and "valleys." Such distributions are natural generalizations of monotone (k=0k=0) and unimodal (k=1k=1) probability distributions, which have been intensively studied in probability theory and statistics. In this paper we consider the problem of \emph{learning} (i.e., performing density estimation of) an unknown kk-modal distribution with respect to the L1L_1 distance. The learning algorithm is given access to independent samples drawn from an unknown kk-modal distribution pp, and it must output a hypothesis distribution p^\widehat{p} such that with high probability the total variation distance between pp and p^\widehat{p} is at most ϵ.\epsilon. Our main goal is to obtain \emph{computationally efficient} algorithms for this problem that use (close to) an information-theoretically optimal number of samples. We give an efficient algorithm for this problem that runs in time poly(k,log⁡(n),1/ϵ)\mathrm{poly}(k,\log(n),1/\epsilon). For k≤O~(log⁡n)k \leq \tilde{O}(\log n), the number of samples used by our algorithm is very close (within an O~(log⁡(1/ϵ))\tilde{O}(\log(1/\epsilon)) factor) to being information-theoretically optimal. Prior to this work computationally efficient algorithms were known only for the cases k=0,1k=0,1 \cite{Birge:87b,Birge:97}. A novel feature of our approach is that our learning algorithm crucially uses a new algorithm for \emph{property testing of probability distributions} as a key subroutine. The learning algorithm uses the property tester to efficiently decompose the kk-modal distribution into kk (near-)monotone distributions, which are easier to learn.

Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.