Papers
Topics
Authors
Recent
Search
2000 character limit reached

Parameter estimation for integer-valued Gibbs distributions

Published 5 Apr 2019 in math.PR, cs.CC, and cs.DM | (1904.03139v6)

Abstract: A central problem in computational statistics is to convert a procedure for sampling combinatorial from an objects into a procedure for counting those objects, and vice versa. Weconsider sampling problems coming from Gibbs distributions, which are probability distributions of the form μ<sup>Ωβ(ω)</sup>e<sup>β</sup>H(ω)\mu<sup>\Omega_\beta(\omega)</sup> \propto e<sup>{\beta</sup> H(\omega)} for β\beta in an interval $[\beta_\min, \beta_\max]$ and H(ω)0[1,n]H( \omega ) \in {0 } \cup [1, n]. The partition function is the normalization factor Z(β)=ωΩe<sup>β</sup>H(ω)Z(\beta)=\sum_{\omega \in\Omega}e<sup>{\beta</sup> H(\omega)}. Two important parameters are the log partition ratio $q = \log \tfrac{Z(\beta_\max)}{Z(\beta_\min)}$ and the vector of counts cx=H<sup>1(x)c_x = |H<sup>{-1}(x)|. Our first result is an algorithm to estimate the counts cxc_x using roughly O~(qϵ<sup>2)\tilde O( \frac{q}{\epsilon<sup>2}) samples for general Gibbs distributions and O~(n<sup>2ϵ<sup>2</sup></sup>)\tilde O( \frac{n<sup>2}{\epsilon<sup>2}</sup></sup> ) samples for integer-valued distributions (ignoring some second-order terms and parameters). We show this is optimal up to logarithmic factors. We illustrate with improved algorithms for counting connected subgraphs and perfect matchings in a graph. We develop a key subroutine for global estimation of the partition function. Specifically, we produce a data structure to estimate Z(β)Z(\beta) for \emph{all} values β\beta, without further samples. Constructing the data structure requires O(qlognϵ<sup>2)O(\frac{q \log n}{\epsilon<sup>2}) samples for general Gibbs distributions and O(n<sup>2</sup>lognϵ<sup>2</sup>+nlogq)O(\frac{n<sup>2</sup> \log n}{\epsilon<sup>2}</sup> + n \log q) samples for integer-valued distributions. This improves over a prior algorithm of Kolmogorov (2018) which computes the single point estimate $Z(\beta_\max)$ using O~(qϵ<sup>2)\tilde O(\frac{q}{\epsilon<sup>2}) samples. We also show that this complexity is optimal as a function of nn and qq up to logarithmic terms.

Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.