Papers
Topics
Authors
Recent
Search
2000 character limit reached

State Complexity of Pattern Matching in Regular Languages

Published 12 Jun 2018 in cs.FL | (1806.04645v2)

Abstract: In a simple pattern matching problem one has a pattern ww and a text tt, which are words over a finite alphabet Σ\Sigma. One may ask whether ww occurs in tt, and if so, where? More generally, we may have a set PP of patterns and a set TT of texts, where PP and TT are regular languages. We are interested whether any word of TT begins with a word of PP, ends with a word of PP, has a word of PP as a factor, or has a word of PP as a subsequence. Thus we are interested in the languages (PΣ<sup>∗)∩</sup>T(P\Sigma<sup>*)\cap</sup> T, (Σ<sup>∗P)∩</sup>T(\Sigma<sup>*P)\cap</sup> T, (Σ<sup>∗</sup>PΣ<sup>∗)∩</sup>T(\Sigma<sup>*</sup> P\Sigma<sup>*)\cap</sup> T, and (Σ<sup>∗</sup>shu⁡P)∩T(\Sigma<sup>*</sup> \mathbin{\operatorname{shu}} P)\cap T, where shu⁡\operatorname{shu} is the shuffle operation. The state complexity κ(L)\kappa(L) of a regular language LL is the number of states in the minimal deterministic finite automaton recognizing LL. We derive the following upper bounds on the state complexities of our pattern-matching languages, where κ(P)≤m\kappa(P)\le m, and κ(T)≤n\kappa(T)\le n: κ((PΣ<sup>∗)∩</sup>T)≤mn\kappa((P\Sigma<sup>*)\cap</sup> T) \le mn; κ((Σ<sup>∗P)∩</sup>T)≤2<sup>m−1n\kappa((\Sigma<sup>*P)\cap</sup> T) \le 2<sup>{m-1}n; κ((Σ<sup><em>PΣ</em>)∩</sup>T)≤(2<sup>m−2+1)n\kappa((\Sigma<sup><em>P\Sigma^</em>)\cap</sup> T) \le (2<sup>{m-2}+1)n; and κ((Σ<sup>∗shu⁡</sup>P)∩T)≤(2<sup>m−2+1)n\kappa((\Sigma<sup>*\mathbin{\operatorname{shu}}</sup> P)\cap T) \le (2<sup>{m-2}+1)n. We prove that these bounds are tight, and that to meet them, the alphabet must have at least two letters in the first three cases, and at least m−1m-1 letters in the last case. We also consider the special case where PP is a single word ww, and obtain the following tight upper bounds: κ((wΣ<sup>∗)∩</sup>Tn)≤m+n−1\kappa((w\Sigma<sup>*)\cap</sup> T_n) \le m+n-1; κ((Σ<sup>∗w)∩</sup>Tn)≤(m−1)n−(m−2)\kappa((\Sigma<sup>*w)\cap</sup> T_n) \le (m-1)n-(m-2); κ((Σ<sup><em>wΣ</em>)∩</sup>Tn)≤(m−1)n\kappa((\Sigma<sup><em>w\Sigma^</em>)\cap</sup> T_n) \le (m-1)n; and κ((Σ<sup>∗shu⁡</sup>w)∩Tn)≤(m−1)n\kappa((\Sigma<sup>*\mathbin{\operatorname{shu}}</sup> w)\cap T_n) \le (m-1)n. For unary languages, we have a tight upper bound of m+n−2m+n-2 in all eight of the aforementioned cases.

Citations (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.