Papers
Topics
Authors
Recent
Search
2000 character limit reached

How good is Good-Turing for Markov samples?

Published 3 Feb 2021 in cs.IT, math.IT, math.ST, stat.ML, and stat.TH | (2102.01938v3)

Abstract: The Good-Turing (GT) estimator for the missing mass (i.e., total probability of missing symbols) in nn samples is the number of symbols that appeared exactly once divided by nn. For i.i.d. samples, the bias and squared-error risk of the GT estimator can be shown to fall as $1/n$ by bounding the expected error uniformly over all symbols. In this work, we study convergence of the GT estimator for missing stationary mass (i.e., total stationary probability of missing symbols) of Markov samples on an alphabet X\mathcal{X} with stationary distribution [πx:x∈X][\pi_x:x \in \mathcal{X}] and transition probability matrix (t.p.m.) PP. This is an important and interesting problem because GT is widely used in applications with temporal dependencies such as LLMs assigning probabilities to word sequences, which are modelled as Markov. We show that convergence of GT depends on convergence of (P<sup>∼</sup>x)<sup>n(P<sup>{\sim</sup> x})<sup>n, where P<sup>∼</sup>xP<sup>{\sim</sup> x} is PP with the xx-th column zeroed out. This, in turn, depends on the Perron eigenvalue λ<sup>∼</sup>x\lambda<sup>{\sim</sup> x} of P<sup>∼</sup>xP<sup>{\sim</sup> x} and its relationship with πx\pi_x uniformly over xx. For randomly generated t.p.ms and t.p.ms derived from New York Times and Charles Dickens corpora, we numerically exhibit such uniform-over-xx relationships between λ<sup>∼</sup>x\lambda<sup>{\sim</sup> x} and πx\pi_x. This supports the observed success of GT in LLMs and practical text data scenarios. For Markov chains with rank-2, diagonalizable t.p.ms having spectral gap β\beta, we show minimax rate upper and lower bounds of 1/(nβ<sup>5)1/(n\beta<sup>5) and 1/(nβ)1/(n\beta), respectively, for the estimation of stationary missing mass. This theoretical result extends the $1/n$ minimax rate for i.i.d. or rank-1 t.p.ms to rank-2 Markov, and is a first such minimax rate result for missing mass of Markov samples.

Citations (1)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.