The mean and variance of - Algorithms Project

Hidden Pattern Statistics
Philippe Flajolet , Yves Guivarc’h ,
Wojciech Szpankowski , and Brigitte Vallée
Algorithms
Project, INRIA-Rocquencourt, 78153 Le Chesnay, France
IRMAR,
de Rennes I, F-35042 Rennes Cedex, France
Dept. ComputerUniversité
Science, Purdue University, W. Lafayette, IN 47907, U.S.A
GREYC, Université de Caen, F-14032 Caen Cedex, France
Abstract. We consider the sequence comparison problem, also known as “hidden pattern” problem, where one searches for a given subsequence in a text (rather
than a string understood as a sequence of consecutive symbols). A characteristic
parameter is the number of occurrences of a given pattern of length
as a
subsequence in a random text of length generated by a memoryless source.
Spacings between letters of the pattern may either be constrained or not in order
to define valid occurrences. We determine the mean and the variance of the number of occurrences, and establish a Gaussian limit law. These results are obtained
via combinatorics on words, formal language techniques, and methods of analytic combinatorics based on generating functions and convergence of moments.
The motivation to study this problem comes from an attempt at finding a reliable
threshold for intrusion detections, from textual data processing applications, and
from molecular biology.
1 Introduction
String matching and sequence comparison are two basic problems of pattern matching known informally as “stringology”. Hereafter, by a string we mean a sequence of
(of length
consecutive symbols. In string matching, given a pattern
) one searches for some/all occurrences of (as a block of consecutive symbols) in
a text
of length . The algorithms by Knuth–Morris–Pratt and Boyer–Moore [7]
provide efficient ways of finding such occurrences. Accordingly, the number of string
occurrences in a random text has been intensively studied over the last two decades, with
significant progress in this area being reported [3, 9, 10, 15–17, 24]. For instance Guibas
and Odlyzko [9, 10] have revealed the fundamental rôle played by autocorrelation vectors and their associated polynomials. Régnier and Szpankowski [16, 17] established
that the number of occurrences of a string is asymptotically normal under a diversity of
models that include Markov chains. Nicodème, Salvy, and Flajolet [15] showed generally that the number of places in a random text at which a ‘motif’ (i.e., a general regular
expression pattern) terminates is asymptotically normally distributed.
In sequence comparisons, we search for a given pattern
in the
text
as a subsequence, that is, we look for indices
such that
,
, ,
. We also say that the word is
“hidden” in the text; thus we call this the hidden pattern problem. For example, date
occurs as a subsequence in the text hidden pattern, in fact four times, but not
"*+(
- ,/.!0 -,213 ' '' - ,546
%$ & $(''')$
!#" "
2
"=998> " =::>@8? ;8 "<
7
even once as a string. We can impose an additional set of constraints on the indices
to record a valid subsequence occurrence: for a given family of integers
(
, possibly
), one should have
. In other words, the
allowed lengths of the “gaps” (
) should be
. With representing a
‘don’t-care-symbol’ (similar to the unix ‘ ’-convention) and the subscript denoting a
strict upper bound on the length of the associated gap, a typical pattern may look like
ab r ac a d a br a;
there, abbreviates
and
is omitted; the meaning is that ‘ab’ should occur first
contiguously, followed by ‘r’ with a gap of
symbols, followed anywhere later in
the text by ‘ac’, etc. The case when all the ’s are infinite is called the unconstrained
problem; when all the ’s are finite, we speak of the constrained problem. The case
where all reduce to 1 gives back classical string matching as a limit case.
I
=:>
=9> A
>ED >H =:>
E" > D F " > F BC" GF "$ = > I
J
I I I I I I I
I@K I
=:> $ L
=:>
Motivations. Our original motivation to study this problem came from intrusion
detection in the area of computer security. The problem is important due to the rise
of attacks on computer systems. There are several approaches to intrusion detections,
but, recently the pattern matching approach has found many advocates, most notably
in [2, 14, 25]. The main idea of this approach is to search in an audit file (the text) for
certain patterns (known also as signatures) representing suspicious activities that might
be indicative of an intrusion by an outsider, or misuse of the system by an insider. The
key to this approach is to recognize that these patterns are subsequences because an
intrusion signature specification requires the possibility of a variable number of intervening events between successive events of the signature. In practice one often needs
to put some additional restrictions on the distance between the symbols in the searched
subsequence, which leads to the constrained version of subsequence pattern matching.
The fundamental question is then: How many occurrences of a signature (subsequence)
constitute a real attack? In other words, how to set a threshold so that we can detect only real intrusions and avoid false alarms? It is clear that random (unpredictable)
events occur and setting the threshold too low will lead to an unrealistic number of false
alarms. On the other hand, setting the threshold too high may result in missing some attacks, which is even more dangerous. This is a fundamental problem that motivated our
studies of hidden pattern statistics. By knowing the most likely number of occurrences
and the probability of deviating from it, we can set a threshold such that with a small
probability we miss real attacks.
Molecular biology provides another important source of applications [18, 23, 24].
As a rule, there, one searches for subsequences, not strings. Examples are in abundance: split genes where exons are interrupted by introns, starting and stopping signal
in genes, tandem repeats in DNA, etc. In general, for gene searching, the constrained
hidden pattern matching (perhaps with an exotic constraint set) is the right approach for
finding meaningful information. The hidden pattern problem can also be viewed as a
close relative of the longest common subsequence (LCS) problem, itself of immediate
relevance to computational biology and still surrounded by mystery [20].
We, computer scientists and mathematicians, are certainly not the first who invented
hidden words and hidden meaning [1]. Rabbi Akiva in the first century A.D. wrote a collection of documents called Maaseh Merkava on secret mysticism and meditations. In
the eleventh century Spanish Solomon Ibn Gabirol called these secret teachings Kab-
3
balah. Kabbalists organized themselves as a secret society dedicated to study of the
ancient wisdom of Torah, looking for mysterious connections and hidden truth, meaning, and words in Kaballah and elsewhere (without computers!). Recent versions of this
activity are knowledge discovery and data mining, bibliographic search, lexicographic
research, textual data processing, or even web site indexing. Public domain utilities like
agrep, grappe, webglimpse (developed by Manber and Wu [26], Kucherov [13],
and others) depend crucially on approximate pattern matching algorithms for subsequence detection. Many interesting algorithms, based on regular expressions and automata, dynamic programming, directed acyclic word graphs, digital tries or suffix trees
have been developed; see [5, 8, 13, 26] for a flavour of the diversity of approaches.
In all of the contexts mentioned above, it is of obvious interest to discern what
constitutes a meaningful observation of pattern occurrences from what is merely a statistically unavoidable phenomenon (noise!). This is precisely the problem addressed
here. We establish hidden pattern statistics—i.e., precise probabilistic information on
number of occurrences of a given pattern as a subsequence in a random text
generated by a memoryless source, this in the most general case (covering the constrained
and unconstrained versions as well as mixed situations). Surprisingly enough and to the
best of our knowledge, there are no results in the literature that address the question at
this level of generality. An immediate consequence of our results is the possibility to
set thresholds at which appearance of a (subsequence) pattern starts being meaningful.
M
be the number of occurrences of a given pattern
as a subseResults. Let
quence in a random text of length generated by a memoryless source (i.e., symbols
are drawn independently). We investigate the general case where we allow some of the
gaps to be restricted, and others to be unbounded. Then the most important parameter
is the quantity defined as the number of unbounded gaps (the number of indices for
which
) plus 1; the product of all the finite constraints
plays also a rôle.
We obtain the mean, the variance, all moments, and finally a central limit law. Precisely,
we prove in Theorem 1 that the number of occurrences has mean and variance given by
=:> PA N
=:>
Q
O
RTS M VUW Y X \Q [B] H 8_^!`ba S M cUWd Be H Xgf
N9Z
H
d H is a computable constant that depends
where [Be
is the probability of , and B]
explicitly (though intricately) on the structure of the pattern and the constraints. Then
we prove the central
moment methods, that is, we show that all centered
F RTS M limit
U Hih YlawXgf 1. byconverge
moments B-M
to the appropriate moments of the Gaussian
distribution (Theorem 2). We stress that, except in the constrained case, the difficulty
of the analysis lies in a nonlinear growth of the mean and the variance so that many
standard approaches to establishing the central limit law tend to fail.
For the unconstrained problem, one has
, and both the mean and the variance
admit pleasantly simple closed forms. For the constrained case, one has
, while the
mean and the variance become of linear growth. To visualize the dependency of
of , we observe that, when all the
equal 1, the problem reduces to traditional
string matching that was extensively studied in the past as witnessed by the (incomplete)
list of references: [3, 9, 10, 15–17, 24]. It is well known that for string matching the
variance coefficient is a function of the so-called autocorrelation of the string. In the
N =:>
N d H
]B 4
general case of hidden pattern matching, the autocorrelation must be replaced by a
more complex quantity that depends on the way pairs of constrained occurrences may
intersect (cf. Theorem 1).
Methodology. The way we approach the probabilistic analysis is through a formal
description of situations of interest by means of regular languages. Basically such a
description of contexts of one, two, or several occurrences gives access to expectation,
variance, and higher moments, respectively. A systematic translation into generating
functions is available by methods of analytic combinatorics deriving from the original Chomsky-Schützenberger theorem. Then, the structure of the implied generating
provides the necessary asymptotic information. In fact,
functions at the pole
there is an important phenomenon of asymptotic simplification where the essentials of
combinatorial-probabilistic features are reflected by the singular forms of generating
functions. For instance, variance coefficients come out naturally from this approach together with, for each case, a suitable notion of correlation; higher moments are seen
to arise from a fundamental singular symmetry of the problem, a fact that eventually
carries with it the possibility of estimating moments. From there Gaussian laws eventually result by basic moment convergence theorems. Perhaps the originality of the
present approach lies in such a joint use of combinatorial-enumerative techniques and
of analytic-probabilistic methods.
jk
2 Framework
+ x''' lnmoqpsr s8 r t88 rVuwv 8 H
Be8 = 7 H
yz {'D~'' | H =
:8
;
7B f
H Be!} \" &$ p9Az" !v $0f''')$
98
:
8
8

€
e
B
"
"
<
"
" H
7 " >ED F " > = >
‚ B/7
7
" ƒ
H ‚ B/7 H
8
8
;
8
7
„
€
…
C
B
"
"
"
.†+ H
,
, 1% 88 , 4zz M B/7
7
M BC7 H ˆ‰:‡ŠŒ‹cŽxw‘ ˆ 8 with ‘ ˆ mo S occurs at position € in U 8 (1)
S ’!U if the property ’ holds, and S ’!U “ otherwise (Iverson’s notation).
where
=>
Blocks
and aggregates. In the general case, the subset ” of indices O for which
=
>
is finite ( $ A ) has cardinality F N with •–N . The two extreme values
of N , namely, N& and N—+ , thus describe the (fully) unconstrained= and the (fully)
>
constrained
respectively. The subset ˜ of indices O for which is unbounded
=(:> ™A ) hasproblem
cardinality N F . It separates the pattern into N independent subpat=:>
terns that are called the blocks and are denoted by 98 t8 . All the possible
X
“inside” #u are finite and form the subconstraint 7šu . In the example described in the
introduction, one has N›(œ and the six blocks are
We fix an alphabet
. The text is
. A particular
matching problem is specified by a pair
: the pattern
is a word
of length ; the constraint
is an element of
.
Positions and occurrences. An -tuple
(
) satisfies the constraint if
, in which case it is called a position.
Let
be the set of all positions subject to the separation constraint , satisfying
furthermore
. An occurrence of pattern
in the text
of length subject
to the constraint is a position
of
for which
,
. For a text
of length , the number of occurrences
(of ) subject to the constraint is then a sum of characteristic variables
5
I I I 8 8; 8 H I # I Ÿž
€TBC" " " u ¢
7
N
€¡ *¢ 8 €¡ -¢ 8 €¡ X ¢ ’ u ¢ £
€) Ÿu
7šu u ¢ £
¤| €
¥›Be€ H
N ¦%§©¨Cªt«E¬9«-­t«¯®±°t«¯®±­t«E²s²9«-³s´:«g³s³:«Eµ´t«Eµ9®s«gª´w¶
satisfies
¸ §©¨C³s´t«g³³w¶E«T¦V· » ¸ §z¨eµ´t«gµ9®;¶E«¦V· C¸ §„ª´ ;
*¸ §ºto²s²9six«¦V·subpositions,
¦V· C¸ §z¨Ctheª:«g¬:constraint
«<­w¶E«P¦V· ¹¸ §©7 ¨*and
®±°t«¯®¯gives
­s¶E«¦V·rise
H
accordingly,
¼!· C¸ §©½ ª:«g­¾Cthe«¼!resulting
· ¹¸ §©½o®±°t«¯aggregate
®±­¾C«¿¼!· *¸ §©¥›½ Be²s€ ²¯¾C«¿is¼šformed
· ¸ §©½ ³s´:with
«g³s³¾C«¿six¼!·blocks,
» ¸ §©½ µ´t«iµ:®E¾C«¼!· À ¸ §©½ ªs´;¾CÁ
Probabilistic model. We consider a memoryless source that emits symbols of the
text independently and denote by Έ ( “ $ Âà $ ) the probability of the symbol ¥„Äl
is drawn according
being emitted. For a given length
, a random text, denoted by
H is defined
to the product
probability
on l . For instance, the pattern probability [Be
H
ÆÅ ,5Ç ÂÉÈŒÊ , a quantity thatH surfaces throughout the analysis. Under this
by [B]
randomness model, the quantity M BC7 becomes a random variable that is itself
H a sum
of correlated random variables ˆ (defined in (1)) for all allowable €ËƂ B/7 .
‘
Generating functions. We shall consider throughout this paper structures superimposed on words. For a class Ì of structures and given a weight function Í (induced by
the probabilities of individual letters), we introduce the generating function
Î Bej HÏ ‡ Î j mo ‡ Í:BCÒ H jŒÓ Ð Ó 8
Ð ‰tÑ
where
ÒxÔ is the number of letters involved in the structure. Then , Ρ S j U/Î BejtheH issizethe Ôtotal
weight of all structures of size . The collection of occurrences
=a b r,
= a c,
= a,
= d a,
=b r,
= a.
of
subject to constraint gives
In the same way, an occurrence
rise to subpositions
, the th term
being an occurrence of
subject to constraint . The th block
is the closed segment whose end points are
the extremal elements of
, and the aggregate of position , denoted by
, is the
collection of these blocks. In the example of the introduction, the position
1
is described by means of regular expressions extended with disjoint unions, and Cartesian products. It is then known that disjoint unions and Cartesian products correspond
respectively to sums and products of generating functions; see [19, 21] for a general
framework. Such correspondences make it possible to translate symbolically combinatorial descriptions into generating function equations and a great use is made of this in
what follows. All the resulting generating functions turn out to be rational, of the form
for some integer
and polynomial , so that
Î Bej H Õ Bg F j H f 5Ö D
Î mo
<× B]j H
×
Ø? “
S j U F H Ö D × Bej H Ö × Bg ºH Ù Û Ú#܁B H<Ý B- j
Ø Z
(2)
3 Mean and Variance Estimates of the Number of Occurrences
Þ
Œá ¨Cß9¶ .
in the series
Mean value analysis. The first moment analysis is easily obtained by describing the
collection of all occurrences in terms of formal languages. Let be the collection of
all occurrences of
as a hidden word. Each occurrence can be viewed as a “context”
1
½
w
ß
9
à
â
¾
Œ
á
¨Cßw¶ represents the coefficient of ßwà
The notation
6
Þ
Þkl!ãåä—p vYälæç . ä—p våä|lËæç 1 ä ä—p f vYälæç 4éèc. ä—ps vYäl!ã (3)
=
=
There, for $ A , l æç denotes= the collection of all words of length strictly less , i.e.,
l æç mê6K ë , æç l , , whereas, for, 6A , l æ K denotes the collection of all finite words,
i.e., l æ
mêl ã ìë , æ K l . The associated generating functions are
í B]j H ÕÛÚºj%Ú#j Ú ''' Ú#j ç f F F j ç 8 í KîB]j H ÛÚ#jïÚ#j Ú ''' F ç
j
j
H
S
U
ˆ
We now weight
H each occurrence by the quantity [Be (ð , so that the generating
RTS U ,
function ܁B]j of Þ coincides with the generating function of‘ the expectations M
Ù Ý X D ó
Ê
H
@
R
S
c
U
܁B]j ÕV‡ ñ M j F j äò ,2Ç Â È Ê-jõôöäò , ó ‰:÷ F F j j ç ô 8 (4)
H the probability of the pattern , one finds from (2) and (4):
and, with [Be
R@S M U S j U ܁Bej H X ò ó = , ôø[B] HGÙ ÛÚŸÜ Ù ÝGÝ NwZ , ‰:÷
with an initial string, then the first letter of the pattern, then a separating string, then the
second letter, etc. The collection is then described by
Variance analysis. For variance and higher moment analysis, it is essential to work
with centred random variables defined as
ù ˆ om ˆ F @R S ˆ U ˆ F [Be H 8ûú /B 7 H mêM B/7 H F R@S M BC7 H U ‡ ù ˆ ‘
‘ ‘
ˆ‰:Š ‹ ŽxH 
H
ú
The second moment of the centred variable
/B 7 equals the variance of M B/7 and
with the centred variables defined above one has
R@S ú B/7 H U ‡ R@S ù ˆ ù ý U ˆü ýb‰:ŠÉ‹VŽx
H
8
þ
There are two kinds of pairs B]€
according as theyù intersect
or not. When € and þ
ù
ˆ
ý
do not intersect, the corresponding random variables
are independent, and
S ù ˆ ù ý U reduces to 0. It isandthus sufficient
the corresponding covariance ð
to consider
intersecting subsets € and þ . Suppose that there exist two occurrences of pattern at
positions € and þ which intersect at ÿ distinct places, the Ø -th intersection point being
the £ Ö -th in the natural ordering of € and the Ö -th in the natural ordering of þ . (This is
.) We then denote by ˆ bý the
only possible if, for all Ø 8 &zØ #ÿ , one has |u &
H
þ
subpattern of that occurs at position €
, and by [Be ˆ bý the probability of this
H
h
@
R
S
U
ˆ‘ ‘ ý equals [Be [B] ˆ bý H H , the expectation
subpattern. Since the expectation
RTS ù ˆ ù ý U RTS ˆ ý U F [Be H involves
a correlation number VBe€ 8;þ
‘
‘
R@S ù ˆ ù ý U z[ Be H cB]€ 8þ H 8 with VBe€ 8;þ H ˆ bý H F (5)
[Be
H
Sù ˆ ù ý U,
In this case, we take the pair of occurrences relative to B]€ 8þ as weighted by ð
and consider the collection Þ of pairs of intersecting occurrences. The associated
7
Ü Bej H coincides with the generating function of the expectations
Ü Bej H b‡ ñ j ‡ ‹ ð S ù ˆ ù ý U b‡ ñ j RTS ú BC7 H U H
We now need to estimate Ü B]j as j
. First, define the aggregate ¥›Be€ 8;þ H to
be the system ofH blocks obtained
H by merging together allH intersectingH blocks of the two
aggregates ¥›Be€ and ¥ÛB þ . The number of blocks Be€ 8;þ of ¥›Be€ 8;þ plays a fundamental rôle here, since it measures the degree of freedom of pairs. Since € and þ intersect,
H
H
H
there exists at least one block of ¥ÛB]€ that intersects a block of ¥›B þ , so that Be€ 8;þ isH
8
þ
8
L
F
. Next, we group the sets € according
at most equal ¢ to N
to the value of B]€ þ
H
and write Þ for the collection of intersecting pairs B]€ 8þ of occurrences for which
we introduceH
Be€ 8;þ H equals L N F  . Since there8þ His a fundamental
H 䕂 9translation
H is full invariance,
C
B
7
a notion of full pairs: a pair B]€
of ‚ tB/7
if the aggregate ¥ÛB]€ 8þ
S U
completely covers¢ the interval 8 . (Clearly,
values of
finite.) Then
¢ , where
¢ is thearesubset
H D theä possible
the collection Þ is isomorphic to B]l ã Xgf
of full
¢
H
;
8
þ
L
F
equals N
 . The generating
function of Þ
is accordingly
pairs such that Be€
D
Ü ¢ B]j H Ù F j Ý Xgf ä ’ ¢ Bej H ¢
’ ¢ H
Here, Bej is the generating function of the collection and from our earlier dis=
H
= b, ‰:÷ = , . Now,
cussion, it is a polynomial of degree at most L B F ¢ Ú , with H
S
U
an easy dominant pole analysis entails that j Ü Ü
S U <¢ BC X-f . This proves that theH
dominant contribution to the variance is given by j *¢ Ü , which is of order ܁BC X-f .
RTS ú U involves the constant ’ B- H that is the total weight of the
Then, the variance
<¢
’ <¢ B]j H is itself the generating function of the collection
collection
;
the
polynomial
<¢ , conceptually an extension of Guibas and Odlyzko’s
polynomial.
i H ,autocorrelation
Since
the
standard
deviation
is
of
an
order,
that
is
smaller
than
the mean,

Ü
C
B
Y
g
X
f
܁BCYX H , concentration of distribution holds, via a well-known argument based on Chebyshev’s inequalities. In summary:
Ï
Theorem
Consider a general constraint 7 and the number of occurrences M
M BC7 H . The1. mean
and variance of M satisfy
R@S M cU [Be H Ù > ó =9> Ý X Ù Ûڟ܁B H Ý 8
NwZ ‰:÷
^!`ba S M cU d B] H Xgf Ù Ûڟ܁B H Ý 8
= > $ A , and the “variance coefficient” d B] H
where ” is the set of O such thatH
involves the autocorrelation YB]
d Be H [L BeF H H Be H with Be H mê ‡ Ù ˆ bý H F Ý (6)
BN Z
2ˆü ý:<‰ 1 . [Be
ð Sù ˆ ù ý U
generating function
, that is,
"!
#
#
#
$
#
'
&%
$
&%
$
)(
'
$
(
$
$
#
$
$
$
$
(
$
+*-,.
$
/$
(
(
0
1
1
1
3254 76
8
H<¢ is the collectionH of all pairs of occurrences H B]€ 8þ H that satisfy three conBe" 8 they are full; BC"*" they are intersecting;
BeH "¹"*" there is a single pair
BC£ 8 þ HH
¢
¢
’
u
Ë6£ 0N for which the £ th block of ¥›Be€ and the th block T of ¥ÛB
H reComputation
of the variance. The computation of the autocorrelation åBe
H
H of
duces to N computations of correlations åBeŸu 8 , relative
to pairs BeŸu 8 H
8
blocks. Note that each correlation of the form YB]#u involves a totally constrained
problem and can be evaluated by dynamic programming. Precisely, one has
Be H Q ‡
Ù £Ú F F„L ݕ٠L N F F £ F Ý YB] u 8 H 8 (7)
N £
u ü X QuQ £ H is the sum of the cB]€ 8þ H taken over all full intersecting pairs B]€ 8þ H
where YB]#u 8 formed with an occurrence € of #u subject to constraint 7šu and an occurrence þ H of subject*¢ to constraint 7 . Let us explain the formula (7) in words: for a pair Be€ 8;þ of the
H
set ¢ , there is a single pair BC£ 8 ¢ of indices with !#£ 8 !zN for which the £ th block
’ u of ¥›Be€ H and the th¢ block¢ T of ¥›B þ H intersect. Then, there exist £xÚ FL blocks
’ u H
before the block ¥ÛB H 8 T and L N F £ F blocks after
have three> ¢ different
¢ BC" it.$ We£ H then
H,
’
,
$
>
T
5
B
O
degrees of freedom: BC" the relative order of blocks
and
blocks
¢
¢
H
H
H
’
,
and similarly the relative order of> blocks BC" 6£ and blocks B5O
H ; BC"*" the
lengths of the blocks (there are Q ¢ possible¢ lengths for the O th block); BC"*"*" finally the
’ u
relative positions of the blocks and T .
In the unconstrained problem, the parameter N equals , and each block #u is
H simplifies to
reduced to the symbol u . Then the “ correlation coefficient” Be
Be H mo ‡ Ù £Ú F„L ݕ٠L F £ F Ý BC£ 8 HGÙ F Ý 8 (8)
F £
ÂÈ
uü £ F H mo S uïz U .
where the “autocorrelation matrix” of pattern is defined by BC£ 8
The set (
ditions:
with
intersect.
8
1
1
9
1
:
1
1
:
:
;
1
<
;
:
9
(
8
8
=
8
?>
8
8
@>A
1
CB
:
1
;
5D
;
B
B
4 Central Limit Laws
M
Our goal is to prove that
appropriately normalized tends to the standard normal
distribution. We consider the following normalized random variable
ú om Xgú f i M F X-f R@iS M U 8
behaves
where N is the number of blocks of the constraint 7 . We shall show that ú
d
asymptotically as a normal variable with mean 0 and standard deviation . By the
classical moment convergence theorem (Theorem 30.2 of [4]) this is established once
are known to converge to the appropriate moments of the standard
all moments of ú
normal distribution. We remind the reader that if is a standard normal variable (i.e.,
a Gaussian distributed variable with mean 0 and standard deviation 1), then for any
? “
integral
RTS U Õ ' ''' B L F H 8 R@S D U (“ (9)
E
0
0
E
E
F
F
G
F
9
£ £î L and £î
L ÛÚ
R@S ú D U bBC  D *  -X f i  H 8 TR S ú UW(d Bg ' ''' B L F HgH Xgf 8 (10)
which implies Gaussian convergence of ú .
Theorem 2. The random variable M asymptotically follows a Central Limit Law:
RTS cU
K M ^!F `ba S M M cU \ L [ K f 1 E = (11)
f
Proof. The proof below is combinatorial; it basically reduces to grouping and enumerRTS u U .
ating adequately the
various combinations of indices in the sum that expresses ú
H
S
U
Once more, ‚ B/7 is formed of all the positions of 8 subject to the constraint 7
H ‚ BC7 H . Then totally distributing the terms in ú u B/7 H yields
and ‚ B/7 ë
RTS ú u U RTS ù ˆ . ''' ù ˆ U ‡
(12)
2ˆ . ü ü ˆ *‰:Š ‹ Žx
H
H   said to be friendly if each € Ö intersects at
An £ -tuple of sets B]€ 98;8 €;u in ‚ u B/7 is
H
least one other € , with
and we let u B/7 be the set of all friendly collections in
ÿ
ì
Ø


H
theH subscript each time
‚ u B/7 . For ‚ u , u , and their derivatives below, we8add
 u  B/7 theH ,
situation is particularized to texts of length . If Be€ 8 € u does not lie in
RTS ù ˆ . ''' ù ˆ U (“ 8 since at least one of the ù ˆ ’s is independent of the other factors
then
ù
R@S ù ˆ U (“ . One can thus restrict attention
in the product and the ˆ ’s have been centred,
to friendly families and get the basic formula
R@S ú u U R@S ù ˆ . ''' ù ˆ U 8
‡
(13)
2ˆ . ü ü ˆ *‰ ‹ Žx
We shall accordingly distinguish two cases based on the parity of ,
, and prove that
0
+H
G
E
IKJ
ML *
NPOPQ
R
9SUT
^]^]^] D
_
a`
V
XWZY
\[
0
D
D
b
b
b
D
^]^]^]
D D /c
D
where the expression involves fewer terms than in (12). From there, we proceed in two
stages. First, restrict attention to friendly families that give rise to the dominant contrib
bution and introduce a suitable subfamily b
; in so doing, moments of odd
ed
order appear to be negligible. Next, for even order , the family b
involves a symmethat corresponds
try and it suffices to consider another smaller subfamily b
b
d
to a “standard” form of occurrence intersection; this last reduction precisely gives rise
to the even Gaussian moments.
Zb
Odd moments. Given
, one defines the aggregate
as the aggregation (in the sense of the variance calculation above) of
. Next, the number of blocks of
is the number of blocks of the aggregate
; if is the total number of intersecting blocks of the aggregate
, the aggregate
has
blocks. Like previously, we
say that the family
of b % is full if the aggregate
completely covers the interval f' . In this case, the length of the aggregate is at most
, and the generating function of full families is a polynomial
of
u 


ã £ u
u  ãu u 
ãiã ã
B]€ 88 € u H Ä  u 
¥ÛB]€;u H 988 H
B]€ :8;8 €u H
¥ÛB]€ €u  98 t8 H
¥ÛB]€ 988 €u H
¥›H Be€ €  u  €;u £:N F Â
98
;
8
B]€ S €;8 u U
£ = B F H Ú0
¥ÛB]€ 8 H € ~ 8'';' 8 ~€ u H
¥ÛB]€
¥›Be€ :8 € :8 €;u H
× ucBej H
10
 u  £ = B F H (Ú =
> ‰:÷ =:>
D Ø
Ö
Ù F Ý ä × tu B]j H 8
 j
u whose block number equals Ø is Ü BC Ö H . This
so that the number of families of
observation proves that the dominant contribution to (13) arises from friendly families
with a maximal block number.
It is clear that h the minimum number of intersecting
 u  equals
equals /£ L , since it coincides exactly with the
blocks of any element of
minimum number of edges of a graph with £ vertices which contains no h isolated vertex.
Then the maximum block number of a friendly family equals £:N F C£ L . In view of
this fact and the remarks above regarding cardinalities, we immediately have
R ú D (Ü  D  Xgf f  D * Xgf i 
degree at most
with g*-,.
. Then, the generating function of
families of b
whose block number equals is of the form
b
b
i
h
i
h
kj
l
km
n
0
oHpm
n
which establishes the limit form of odd moments in (10).
Even moments. We are thus left with estimating the even moments. The dominant
term is relative to friendly families of b with an intersecting block number equal to
, whose set we denote by b
. In such a family, each subset intersects one and only
one other subset _ . Furthermore, if the blocks of q are denoted by q r
,
Cs
r
u
r
t
there exists only one block
of
and only one block _
that contains the
and v
for all
points of
_ . This defines an involution v such that v
pairs of indices
for which
and _ intersect. Furthermore, given the symmetry
xy xy xw it suffices to restrict attention to friendly
relation
xw
 
 
€Ö
’ ¢ 8 ! zN
¢
¢
’ Ö Û¥ B]€ Ö H
’ H
H 0Ø
€Ö € 8 H
¿
¹
B
Ø
(
ÿ
¿
/
B
ÿ
RTS ù ˆ . ''BC'ÿ ù ؈ 1 U R@S ù ˆ €. Ö ''' ù ˆ € 1 U
  for which the involution is the standard one with cycles B- 8±L H , B ¡8 H ,
families of
  , the pairs that intersect
ã
etc; for such “standard”
families H whose set is denoted by
H
±
ã
ã
8 € . Since
are thus Be€ 8 € , . . . , Be€ the set of involutions of L elements has
H
f
cardinality ' &' ''' B L F 8 the equality
(14)
‡ 1 ‹ R@S ù ˆ . ''' ù ˆ 1 U ˇ 1 ‹ RTS ù ˆ . ''' ù ˆ 1 U 8
entails that we can work now solely with standard families.
H Xgf f ä ¢ ä
The class of occurrences relative¢ to standard families is l ã ä Bel ã
l ã 8 and involves the collection of all¢ full friendly L -tuples of occurrences<¢with
a number of blocks equal to . Since is exactly a shuffle of copies of (as
introduced in the study of the variance), the associated generating function is
D
Ù F Ý Xgf B L wN F H Zò ’L <¢F Bej H H ô 8
j
BN Z
*¢
’ H
where Bej is the already introduced autocorrelation polynomial. Upon taking coefficients, we obtain the estimate
(15)
‡ 1 ‹ R@S ù ˆ . ''' ù ˆ 1 UW  Xgf  d ã
€
b
¥ÛB]€ H
G z
v
|
b
G u}
{
~w
c  ~w +|
~w
c  x w (
(
c  x w €(
(
~w
11
In view of the formulæ (12), (13), (14), and (15) above, this yields the estimate of even
moments and leads to the second relation of (10). (Note that the even Gaussian moments
eventually come out of the number of involutions, which corresponds to a fundamental
symmetry present in the problem.) This completes the proof of Theorem 2.
5 Conclusion
As a test case, we took the full text of Hamlet where all nonalphabetic characters are
suppressed. This gives us a (rather unpoetical looking) text that has one long line with
}M alphabetical characters: “who s there nay answer me
30,316 words and
stand and unfold yourself long live the king bernardo he you come most carefully upon
your hour [. . . ]”. E The pattern is “The law is Gaussian” [
thelawisgaussian] and
its mirror image , corresponding to
. Based on the empirical distribution
of
‚
letter frequencies in the text, we anticipate the pattern to‚ appear GMG f‚
times as
a subsequence, while the observed counts are G }
and G3ƒMƒ
, a deviation
of less than 4% from what is expected. Similarly, if we bound the separation distance
between any two letters by , analysis predicts that the pattern might start occurring
near
, while its presence
is unlikely for smaller values,
. In fact, starts
E
z while
G —a deviation of some 30–40% from what
occurring at
starts at
the model predicts. Here is a table of observed versus predicted values when varies:
\ L “ 8 “
œ
= ös“
=
=
œ s“ ƒ
= “ s“ “
= $ s“
=
§ thelawisgaussian § naissuagsiwaleht
¨
%
¶
Expected
Occurred ( )
Occurred ( )
E
„
13
14
20
50
‰
7…
9.195E+01
2.794E+02
5.886E+04
5.482E+10
1.330E+48
†
0
693
124,499
76,146,232,395
1.36554E+48
†ˆ‡
…
0.00
2.47
2.11
1.38
1.03
†
18
371
41,066
48,386,404,680
1.38807E+48
†ˆ‡
…
0.19
1.32
0.69
0.88
1.04
This (together with many other experiments) shows a fair fit between the theoretical
model and the observed data even though the text chosen is far from being “random”.
Extensions. For the constrained case where all the distances are finite, based on
finite state models and the de Bruijn graph, it is possible to obtain local limit laws (i.e.,
a direct estimation of probability densities),
a characterization of the speed of conver0
gence to the asymptotic limit (it is
), as well as large deviation estimates (that are
exponentially small); see the full paper. For the unconstrained case, the corresponding
problems appear to be related to products of random matrices and to the difficult case
of random walks on nilpotent Lie groups; see Guivarc’h’s paper [11] for context and
references. Finally, preliminary investigations indicate that the methods developed here
apply to Markovian sources and more generally to all dynamical sources in the sense of
Vallée [6, 22].
éf i
Acknowledgments. We thank M. Atallah (Purdue U.) for introducing us to the intrusion detection
problem that motivated this study. This research was supported in part by sponsors of CERIAS
at Purdue under contract 1419991431A, by the A LCOM -FT Project (# IST-1999-14186) of the
European Union, and by NSF Grant C-CR 9804760.
12
References
1. A. Aczel, The Mystery of the Aleph. Mathematics, the Kabbalah, and the Search for Infinity,
Four Walls Eight Windows, New York, 2000.
2. A. Apostolico and M. Atallah, Compact Recognizers of Episode Sequences, Submitted to
Information and Computation.
3. E. Bender and F. Kochman, The Distribution of Subword Counts is Usually Normal, European Journal of Combinatorics, 14, 265-275, 1993.
4. P. Billingsley, Probability and Measure, Second Edition, John Wiley & Sons, New York,
1986.
5. L. Boasson, P. Cegielski, I. Guessarian, and Yuri Matiyasevich, Window-Accumulated Subsequence Matching Problem is Linear, In Proceedings of the Eighteenth ACM SIGMODSIGACT-SIGART Symposium on Principles of Database Systems: PODS 1999, ACM Press,
327–336, 1999.
6. J. Clément, P. Flajolet, and B. Vallée, Dynamical Sources in Information Theory: A General
Analysis of Trie Structures, Algorithmica, 29, 307–369, 2001.
7. M. Crochemore and W. Rytter, Text Algorithms, Oxford University Press, New York, 1994.
8. G. Das, R. Fleischer, L. G asieniec, D. Gunopulos, and J. Kärkkäinen, Episode Matching, In
Combinatorial Pattern Matching, 8th Annual Symposium, Lecture Notes in Computer Science vol. 1264, 12–27, 1997.
9. L. Guibas and A. M. Odlyzko, Periods in Strings, J. Combinatorial Theory Ser. A, 30, 19–43,
1981.
10. L. Guibas and A. M. Odlyzko, String Overlaps, Pattern Matching, and Nontransitive Games,
J. Combinatorial Theory Ser. A, 30, 183–208, 1981.
11. Y. Guivarc’h, Marches aléatoires sur les groupes, Fascicule de probabilités, Publ. Inst. Rech.
Math. Rennes, 2000.
12. D. E. Knuth, The Art of Computer Programming, Fundamental Algorithms, Vol. 1, Third
Edition, Addison-Wesley, Reading, MA, 1997.
13. G. Kucherov and M. Rusinowitch, Matching a Set of Strings with Variable Length Don’t
Cares, Theoretical Computer Science 178, 129–154, 1997.
14. S. Kumar and E.H. Spafford, A Pattern-Matching Model for Intrusion Detection, Proceedings of the National Computer Security Conference, 11–21, 1994.
15. P. Nicodème, B. Salvy, and P. Flajolet, Motif Statistics, European Symposium on Algorithms,
Lecture Notes in Computer Science, No. 1643, 194–211, 1999.
16. M. Régnier and W. Szpankowski, On the Approximate Pattern Occurrences in a Text, Proc.
Compression and Complexity of SEQUENCE’97, IEEE Computer Society, 253–264, Positano, 1997.
17. M. Régnier and W. Szpankowski, On Pattern Frequency Occurrences in a Markovian Sequence, Algorithmica, 22, 631-649, 1998.
18. I. Rigoutsos, A. Floratos, L. Parida, Y. Gao and D. Platt, The Emergence of Pattern Discovery
Techniques in Computational Biology, Metabolic Engineering, 2, 159-177, 2000.
19. R. Sedgewick and P. Flajolet, An Introduction to the Analysis of Algorithms, Addison-Wesley,
Reading, MA, 1995.
20. J. M. Steele, Probability Theory and Combinatorial Optimization, SIAM, Philadelphia,
1997.
21. W. Szpankowski, Average Case Analysis of Algorithms on Sequences, John Wiley & Sons,
New York, 2001.
22. B. Vallée, Dynamical Sources in Information Theory: Fundamental Intervals and Word Prefixes, Algorithmica, 29, 262–306, 2001.
23. A. Vanet, L. Marsan, and M.-F. Sagot, Promoter sequences and algorithmical methods for
identifying them, Res. Microbiol., 150, 779-799, 1999.
24. M. Waterman, Introduction to Computational Biology, Chapman and Hall, London, 1995.
25. A. Wespi, H. Debar, M. Dacier, and M. Nassehi, Fixed vs. Variable-Length Patterns For
Detecting Suspicious Process Behavior, J. Computer Security, 8, 159-181, 2000.
26. S. Wu and U. Manber, Fast Text Searching Allowing Errors, Comm. ACM, 35:10, 83–991,
1995.