Papers
Topics
Authors
Recent
Search
2000 character limit reached

Testing with Non-identically Distributed Samples

Published 19 Nov 2023 in cs.DS, cs.IT, cs.LG, math.IT, and stat.ML | (2311.11194v1)

Abstract: We examine the extent to which sublinear-sample property testing and estimation applies to settings where samples are independently but not identically distributed. Specifically, we consider the following distributional property testing framework: Suppose there is a set of distributions over a discrete support of size kk, p<em>1,p2,…,pT\textbf{p}<em>1, \textbf{p}_2,\ldots,\textbf{p}_T, and we obtain cc independent draws from each distribution. Suppose the goal is to learn or test a property of the average distribution, p</em>avg\textbf{p}</em>{\mathrm{avg}}. This setup models a number of important practical settings where the individual distributions correspond to heterogeneous entities -- either individuals, chronologically distinct time periods, spatially separated data sources, etc. From a learning standpoint, even with c=1c=1 samples from each distribution, Θ(k/ε<sup>2)\Theta(k/\varepsilon<sup>2) samples are necessary and sufficient to learn p<em>avg\textbf{p}<em>{\mathrm{avg}} to within error ε\varepsilon in TV distance. To test uniformity or identity -- distinguishing the case that p</em>avg\textbf{p}</em>{\mathrm{avg}} is equal to some reference distribution, versus has ℓ1\ell_1 distance at least ε\varepsilon from the reference distribution, we show that a linear number of samples in kk is necessary given c=1c=1 samples from each distribution. In contrast, for c≥2c \ge 2, we recover the usual sublinear sample testing of the i.i.d. setting: we show that O(k/ε<sup>2</sup>+1/ε<sup>4)O(\sqrt{k}/\varepsilon<sup>2</sup> + 1/\varepsilon<sup>4) samples are sufficient, matching the optimal sample complexity in the i.i.d. case in the regime where ε≥k<sup>−1/4\varepsilon \ge k<sup>{-1/4}. Additionally, we show that in the c=2c=2 case, there is a constant $\rho &gt; 0$ such that even in the linear regime with ρk\rho k samples, no tester that considers the multiset of samples (ignoring which samples were drawn from the same pi\textbf{p}_i) can perform uniformity testing.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.