Papers
Topics
Authors
Recent
Search
2000 character limit reached

Near-Optimal Learning of Tree-Structured Distributions by Chow-Liu

Published 9 Nov 2020 in cs.DS, cs.IT, cs.LG, and math.IT | (2011.04144v2)

Abstract: We provide finite sample guarantees for the classical Chow-Liu algorithm (IEEE Trans.~Inform.~Theory, 1968) to learn a tree-structured graphical model of a distribution. For a distribution PP on Σ<sup>n\Sigma<sup>n and a tree TT on nn nodes, we say TT is an ε\varepsilon-approximate tree for PP if there is a TT-structured distribution QQ such that D(P  ∣∣  Q)D(P\;||\;Q) is at most ε\varepsilon more than the best possible tree-structured distribution for PP. We show that if PP itself is tree-structured, then the Chow-Liu algorithm with the plug-in estimator for mutual information with O~(∣Σ∣<sup>3</sup>nε<sup>−1)\widetilde{O}(|\Sigma|<sup>3</sup> n\varepsilon<sup>{-1}) i.i.d.~samples outputs an ε\varepsilon-approximate tree for PP with constant probability. In contrast, for a general PP (which may not be tree-structured), Ω(n<sup>2ε<sup>−2)\Omega(n<sup>2\varepsilon<sup>{-2}) samples are necessary to find an ε\varepsilon-approximate tree. Our upper bound is based on a new conditional independence tester that addresses an open problem posed by Canonne, Diakonikolas, Kane, and Stewart~(STOC, 2018): we prove that for three random variables X,Y,ZX,Y,Z each over Σ\Sigma, testing if I(X;Y∣Z)I(X; Y \mid Z) is $0$ or ≥ε\geq \varepsilon is possible with O~(∣Σ∣<sup>3/ε)\widetilde{O}(|\Sigma|<sup>3/\varepsilon) samples. Finally, we show that for a specific tree TT, with O~(∣Σ∣<sup>2nε<sup>−1)\widetilde{O} (|\Sigma|<sup>2n\varepsilon<sup>{-1}) samples from a distribution PP over Σ<sup>n\Sigma<sup>n, one can efficiently learn the closest TT-structured distribution in KL divergence by applying the add-1 estimator at each node.

Citations (22)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.