Papers
Topics
Authors
Recent
Search
2000 character limit reached

Efficient average-case population recovery in the presence of insertions and deletions

Published 12 Jul 2019 in cs.DS, cs.IT, cs.LG, and math.IT | (1907.05964v1)

Abstract: Several recent works have considered the \emph{trace reconstruction problem}, in which an unknown source string x0,1<sup>nx\in{0,1}<sup>n is transmitted through a probabilistic channel which may randomly delete coordinates or insert random bits, resulting in a \emph{trace} of xx. The goal is to reconstruct the original string~xx from independent traces of xx. While the best algorithms known for worst-case strings use exp(O(n<sup>1/3))\exp(O(n<sup>{1/3})) traces \cite{DOS17,NazarovPeres17}, highly efficient algorithms are known \cite{PZ17,HPP18} for the \emph{average-case} version, in which xx is uniformly random. We consider a generalization of this average-case trace reconstruction problem, which we call \emph{average-case population recovery in the presence of insertions and deletions}. In this problem, there is an unknown distribution D\cal{D} over ss unknown source strings x<sup>1,,x<sup>s</sup></sup>0,1<sup>nx<sup>1,\dots,x<sup>s</sup></sup> \in {0,1}<sup>n, and each sample is independently generated by drawing some x<sup>ix<sup>i from D\cal{D} and returning an independent trace of x<sup>ix<sup>i. Building on \cite{PZ17} and \cite{HPP18}, we give an efficient algorithm for this problem. For any support size sexp(Θ(n<sup>1/3))s \leq \smash{\exp(\Theta(n<sup>{1/3}))}, for a $1-o(1)$ fraction of all ss-element support sets x<sup>1,,x<sup>s</sup></sup>0,1<sup>n{x<sup>1,\dots,x<sup>s}</sup></sup> \subset {0,1}<sup>n, for every distribution D\cal{D} supported on x<sup>1,,x<sup>s{x<sup>1,\dots,x<sup>s}, our algorithm efficiently recovers D{\cal D} up to total variation distance ϵ\epsilon with high probability, given access to independent traces of independent draws from D\cal{D}. The algorithm runs in time poly(n,s,1/ϵ)(n,s,1/\epsilon) and its sample complexity is poly(s,1/ϵ,exp(log<sup>1/3n)).(s,1/\epsilon,\exp(\log<sup>{1/3}n)). This polynomial dependence on the support size ss is in sharp contrast with the \emph{worst-case} version (when x<sup>1,,x<sup>sx<sup>1,\dots,x<sup>s may be any strings in 0,1<sup>n{0,1}<sup>n), in which the sample complexity of the most efficient known algorithm \cite{BCFSS19} is doubly exponential in ss.

Citations (20)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.