A Fully Sequential Elimination Procedure for Indifference

Submitted to Operations Research
manuscript OPRE-2011-09-480
A Fully Sequential Elimination Procedure for
Indi↵erence-Zone Ranking and Selection with
Tight Bounds on Probability of Correct Selection
Peter I. Frazier
School of Operations Research and Information Engineering, Cornell University, Ithaca, NY 14853, [email protected],
http://people.orie.cornell.edu/pfrazier/
We consider the indi↵erence-zone (IZ) formulation of the ranking and selection problem with independent
normal samples. In this problem, we must use stochastic simulation to select the best among several noisy
simulated systems, with a statistical guarantee on solution quality. Existing IZ procedures sample excessively
in problems with many alternatives, in part because loose bounds on probability of correct selection lead
them to deliver solution quality much higher than requested. Consequently, existing IZ procedures are
seldom considered practical for problems with more than a few hundred alternatives. To overcome this, we
present a new sequential elimination IZ procedure, called BIZ (Bayes-inspired Indi↵erence Zone), whose
lower bound on worst-case probability of correct selection in the preference zone is tight in continuous
time, and nearly tight in discrete time. To the author’s knowledge, this is the first sequential elimination
procedure with tight bounds on worst-case preference-zone probability of correct selection for more than
two alternatives. Theoretical results for the discrete-time case assume that variances are known and have
an integer multiple structure, but the BIZ procedure itself can be used when these assumptions are not
met. In numerical experiments, the sampling e↵ort used by BIZ is significantly smaller than that of another
leading IZ procedure, the KN procedure of Kim and Nelson (2001), especially on the largest problems tested
(214 = 16, 384 alternatives).
1. Introduction
In the use of simulation, one commonly encounters the problem of selecting the best among several
simulated systems, e.g., selecting the method for operating a supply chain with minimum average
cost, or selecting the configuration of an assembly line with maximum throughput. The higher-level
problem of deciding how many simulation samples to take from each system to best support this
selection of the best is called the ranking and selection (R&S) problem. Doing well in R&S requires
balancing the amount of time spent sampling against the quality of the ultimate selection.
1
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
2
We consider the indi↵erence-zone (IZ) formulation of the R&S problem, in which we wish to
correctly select the best alternative with a probability exceeding a user-specified target, whenever
the best alternative is sufficiently separated from the others. A sampling procedure having this
property is said to satisfy the IZ guarantee. This problem has a rich history, dating to the seminal
work Bechhofer (1954), with early work summarized in the monograph Bechhofer et al. (1968).
Research in the area has been active since that time (see, e.g., Paulson (1964), Fabian (1974),
Rinott (1978), Hartmann (1988), Paulson (1994), Nelson et al. (2001), Goldsman et al. (2002),
Hong (2006), Andradóttir and Kim (2010)). This large body of research is summarized in Bechhofer
et al. (1995), and more recent work is reviewed in Swisher et al. (2003), Kim and Nelson (2006,
2007).
The goal in designing IZ sampling procedures is to take as few samples as possible while still
satisfying the IZ guarantee. While early IZ procedures introduced in Bechhofer (1954), Paulson
(1964), Fabian (1974), Rinott (1978), Hartmann (1988, 1991), Paulson (1994) satisfy the IZ guarantee, they provide a probability of correct selection (PCS) much larger than the user-specified
target probability. This is due in part to their use of the Bonferonni inequality, which leads to loose
theoretical PCS bounds, and is problematic because it leads them to sample more than necessary.
This issue becomes more severe as the number of alternatives grows. More recent procedures developed in Kim and Nelson (2001), Goldsman et al. (2002), Hong (2006) have better performance, but
these procedures continue to use the Bonferonni inequality, causing them to be overly conservative
in large problems, in the sense that they sample more than necessary and over-deliver on PCS
targets (Branke et al. 2007). Recent procedures in Kim and Dieker (2011), Dieker and Kim (2012)
avoid the Bonferonni inequality when comparing groups of three alternatives, but again requires
the Bonferonni inequality for more than three alternatives.
In this paper, we develop the Bayes-inspired IZ (BIZ) procedure, a fully sequential elimination
procedure that satisfies the IZ guarantee (given assumptions on the sampling variances), is less
conservative than existing IZ procedures, and samples less as a consequence. This procedure does
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
3
not use the Bonferroni inequality, instead using a novel symmetry based in Bayesian analysis. We
assume independent normal samples, and present versions for both continuous and discrete time.
The continuous-time BIZ procedure has a tight bound on worst-case preference-zone PCS: the
PCS of the least-favorable configuration is exactly equal to the target probability. This is the first
sequential elimination procedure with this property for more than two alternatives. The worst-case
preference-zone PCS of the discrete-time BIZ procedure is shown in numerical experiments to be
extremely close to the target PCS, even for as many as 214 = 16, 384 alternatives. The discrete-time
BIZ procedure generalizes the non-elimination procedure PB⇤ introduced by Bechhofer et al. (1968),
and the tight worst-case preference-zone PCS bound presented here also applies to a continuoustime version of PB⇤ .
Our theoretical results (the IZ guarantee, and tightness of the worst-case preference-zone PCS
bound) assume that the sampling variances are known, and are either common across alternatives,
or are integer multiples of a common divisor. The BIZ procedure itself, however, allows both
known and unknown sampling variances with arbitrary values, and numerical results suggest that
the procedure’s performance is robust to deviations of the sampling variances from the structure
assumed by the theoretical results.
Although our bound on worst-case preference-zone PCS is tight in continuous time, and nearly
tight in discrete time, BIZ’s PCS under configurations that are not least favorable can be strictly
larger than the target. Thus, in these other configurations, BIZ also over-delivers on PCS. Furthermore, Wang and Kim (2012) shows that, for a variant of Paulson’s procedure, the contribution to
over-delivery from the Bonferonni inequality is smaller than from the requirement that PCS be no
smaller than the target for slippage configurations in the preference zone. However, violating this
slippage configuration requirement would violate the IZ guarantee itself, while our results show
that the Bonferonni inequality can be avoided while retaining the IZ guarantee.
Numerical experiments demonstrate that, across a variety of configurations, BIZ’s over-delivery
is much less than that of a leading IZ procedure, the KN procedure of Kim and Nelson (2001). The
KN family of procedures “might be considered state-of-the-art for IZ R&S” (Branke et al. 2007),
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
4
and has been shown to be highly efficient compared to other existing IZ procedures (Malone et al.
2005, Wang and Kim 2012). Thanks to reduced over-delivery, BIZ requires fewer samples than KN
on a variety of problems.
Although the PCS bounds presented in this paper are non-Bayesian, the BIZ procedure is derived
using a Bayesian approach. This derivation employs a Bayesian prior concentrated on slippage
configurations. The proof techniques are reminiscent of results on the relationship between minimax
and Bayesian analysis from decision theory (see, e.g., Berger (1985)). Thus, this work connects the
IZ with the Bayesian formulation of R&S (see, e.g., Gupta and Miescke (1996), Chick and Inoue
(2001), Frazier et al. (2008, 2009), Frazier and Powell (2008), Chick et al. (2010)).
We begin in Section 2 by formally stating the IZ formulation of the R&S problem. We then introduce the BIZ procedure for discrete time in Section 3, first assuming a known common sampling
variance in Section 3.1, and then the allowing sampling variances to be heterogeneous and unknown
in Section 3.2. Our theoretical results, first that the BIZ procedure satisfies the IZ guarantee when
the variance is known and is either common across alternatives or has an integer multiple structure,
and second that it has tight worst-case preference-zone PCS bounds in continuous time, are given
in Section 4. To support this analysis, a continuous-time generalization of the discrete-time BIZ
procedure is also given. Section 5 gives numerical results, including a comparison with the KN
procedure of Kim and Nelson (2001) and the PB⇤ procedure of Bechhofer et al. (1968).
2. Indi↵erence-Zone Ranking and Selection
We have k alternative simulated systems, among which we would like to select the best. Samples
from system x 2 {1, . . . , k } are normally distributed and independent, over time and across alternatives. Let µx and
and
= ( 1, . . . ,
pair µ,
k)
2
x
be the mean and variance of this sampling distribution. Let µ = (µ1 , . . . , µk )
be the corresponding vectors of sampling means and variances. Together, the
are referred to as a system configuration. Our goal is to observe samples sequentially over
time to find which alternative is the best, in the sense of having the largest µx .
Let t = 0, 1, 2, . . . index time, and let Ytx be the sum of the samples observed from alternative x
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
5
2
x)
by time t, so that Ytx is a discrete-time random walk with N (µx ,
increments and Y0x = 0. For
any set A ✓ {1, . . . , k } let YtA be the vector (Ytx : x 2 A), and Yt be the vector Yt = (Yt1 , . . . , Ytk ).
Any R&S procedure observes samples over time, either adaptively or deterministically, and either
choosing to sample all of the alternatives at each time, or only a subset. Based on these samples,
the procedure eventually stops sampling and selects an alternative as its estimate of the best. Call
the selected alternative x̂. The goal in designing an R&S procedure is to take as few samples as
possible, while still accurately selecting the best alternative.
We now define the indi↵erence-zone guarantee, which is a statistical guarantee on the quality of
the solution produced by an R&S procedure. First, we define the probability of correct selection as
PCS(µ, ) = Pµ,
⇢
x̂ 2 arg max µx ,
x
where Pµ, is the probability measure under which samples from system x have mean µx and
variance
2
x.
In the common-variance case, when
2
x
=
2
for all x, we write PCS(µ, ) in place of
PCS(µ, ) and Pµ, in place of Pµ, .
Then, we define the preference zone (PZ), parameterized by
PZ( ) = µ 2 Rk : µ[k]
where µ[k]
µ[k
1]
...
µ[k
1]
> 0, to be the set
,
µ[1] are the sorted components of µ. This is the set of system configura-
tions under which the best alternative is better than the second best by at least . The complement
of the preference zone is called the indi↵erence zone, and is the set of system configurations in
which we are indi↵erent between the best and second best alternatives.
Then, a procedure meets the indi↵erence-zone (IZ) guarantee at P ⇤ 2 (1/k, 1) and
PCS(µ, )
P⇤
> 0 if
for all µ 2 PZ( ).
We assume that P ⇤ > 1/k because IZ guarantees for smaller values of P ⇤ can be met by choosing
x̂ uniformly at random from among {1, . . . , k } without observing any samples. In this definition,
whether a procedure meets the IZ guarantee depends upon , although procedures that satisfy the
IZ guarantee are usually designed to do so for all
2 Rk++ .
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
6
3. The Bayes-inspired IZ (BIZ) Procedure
In this section we define the Bayes-inspired IZ procedure (BIZ), and summarize the theoretical
results shown later in Section 4. This procedure is developed using a Bayesian motivation, although
the PCS bounds and IZ guarantee that we show are non-Bayesian. We first define a version for
known common variance in Section 3.1, and then generalize to unknown and/or heterogeneous
variances in Section 3.2. The BIZ procedure as described in this section operates in discrete time.
In support of theoretical analysis, Section 4 provides generalizations of this discrete-time procedure
that may operate in discrete or continuous time.
3.1. The BIZ Procedure with Known Common Variance
We first define the Bayes-inspired indi↵erence zone (BIZ) procedure for the case of common known
variance, when
2
x
2
are all equal to the constant
. In Section 4, we show that this procedure
satisfies the IZ guarantee, with tight bounds on worst-case preference-zone PCS in continuous time.
Below, in Section 3.2, we generalize to heterogeneous and/or unknown sampling variances.
BIZ is an elimination procedure. It maintains a list of alternatives that are in contention, and at
each point in time, it takes one sample from each alternative in this set. Initially, all alternatives
are in contention, and over time, as samples are observed, alternatives are eliminated. Once an
alternative is eliminated, it may not come back into contention, and will not be considered for
selection when sampling stops. When all but one alternative has been eliminated, this remaining
alternative is selected as the best. In contrast with non-elimination procedures, which sample every
alternative at every time, elimination procedures may eliminate bad alternatives quickly to reduce
sampling e↵ort.
BIZ is parameterized by the values P ⇤ 2 (1/k, 1) and
> 0 for which we desire an IZ guarantee,
and a parameter c satisfying
c 2 [0, 1
(P ⇤ ) k
c = 0 if k = 2.
1
1
]
if k > 2,
(1)
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
7
The parameter c determines how aggressively we eliminate alternatives, and its choice is discussed
below in Section 3.3. We recommend setting it to its maximum value, 1
2
x
also depends on the sampling variances
=
2
(P ⇤ ) k
1
1
. The procedure
.
For each t, x 2 {1, . . . , k }, and subset A ✓ {1, . . . , k }, we define a function
qtx (A) = exp
✓
2
Ytx
◆
X
exp
x0 2A
✓
2
◆
Ytx0 .
(2)
In Section 4, this expression is shown to be equal to a Bayesian posterior probability that alternative
x is the best, given Yt , and given that the best is in the subset of alternatives A.
The BIZ procedure for known common sampling variance is then defined by Alg. 1.
Algorithm 1 BIZ for known common sampling variance, in discrete time
Require: c 2 [0, 1
(P ⇤ ) k
1
{1, . . . , k }, t
1
],
> 0, P ⇤ 2 (1/k, 1), common sampling variance
1:
Let A
0, P
2:
while maxx2A qtx (A) < P do
3:
while minx2A qtx (A)  c do
4:
Let x 2 arg minx2A qtx (A).
5:
Let P
6:
Remove x from A.
P/(1
P ⇤ , Y0x
2
> 0.
0 for each x.
qtx (A)).
7:
end while
8:
Sample from each x 2 A and add this sample to Ytx to obtain Yt+1,x . Then increment t.
9:
10:
end while
Select x̂ 2 arg maxx2A Ytx as our estimate of the best.
The set A is the set of alternatives in contention, and is initially set to contain all of the
alternatives in Step 1. Alternatives can be eliminated either in Step 6, or by the final selection
in Step 10, which e↵ectively eliminates all remaining alternatives except the one selected. These
eliminations are performed based on the current value of (qtx (A) : x 2 A), and an adaptively updated
8
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
threshold P . The threshold P , which is initially set to P ⇤ , can be interpreted as a Bayesian posterior
probability of selecting the best that we must achieve to stop sampling.
Motivation: We motivate the BIZ procedure as follows. First, consider elimination resulting
from exiting the outer “while” loop in Step 2 and going to Step 10. Recall that qtx (A) can be
interpreted, in a Bayesian setting, as the posterior probability that x is the best (given that the best
is in our contention set). The quantity P is a threshold on the probability of correctly selecting the
best that we must achieve to stop sampling. In Step 2, if an alternative exceeds the this threshold
P , then we exit the loop and select it as best in Step 10.
Now, consider elimination resulting from entering the inner while loop in Step 3 and going to
Step 7. The quantity minx2A qtx (A) is this posterior probability for the alternative in contention
that is least likely to be the best. The inner while loop, Step 3, checks whether this minimal
posterior probability is below the threshold c, and if it is, eliminates this alternative by removing
it from A in steps 4 to 7. In addition to removing this alternative from A, the threshold P is
increased, to account for the fact that we may have incorrectly removed the best alternative from
A, and should strengthen the criteria that we must meet in Step 2 to stop.
The behavior of this algorithm is illustrated in Figure 1. The example illustrated has k = 4
alternatives. The lower threshold c is plotted as a horizontal line, and the upper threshold P is
plotted as a line with jumps. The posterior probability qtx (A) for each alternative x is also plotted
versus time t. The figure uses the additional notation ⌧n to indicate the time at which the nth
elimination occurs, and Zn to indicate the alternative eliminated.
Starting from time t = 0, the contention set A contains all 4 alternatives and we plot qtx (A) for
each. At time ⌧1 , minx2A qtx (A) hits the lower threshold c and the alternative Z1 achieving this
minimum is eliminated. The values qtx (A) jump at this time, as an alternative is removed from A.
This jump is small for the two alternatives with small posterior probabilities qtx (A), but is larger
for the one alternative with a higher value. The threshold P also jumps.
Moving forward from time ⌧1 , three alternatives remain in A, and we plot qtx (A) for each. A
second alternative is eliminated at time ⌧2 when its posterior probability qtx (A) hits threshold c.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
9
Final selection
1
Pn
0.8
0.6
qtx(An)
0.4
0.2
c
0
0
Figure 1
time2000
(t)
8000
!6000
Z2)
2 (eliminate
!1 (eliminate
time (t) Z1)
4000
10000
Illustration of the BIZ procedure with k = 4 alternatives. The BIZ procedure follows the posterior
probabilities qtx (A) over time, eliminating alternatives as they hit the lower threshold. Each time an
alternative is eliminated, the upper threshold jumps upward. Eventually, an alternative reaches the
upper threshold and is selected as the best.
(At the same time, the largest posterior probability comes very close to the upper threshold, but
does not hit it.) After time ⌧2 , posterior probabilities qtx (A) are plotted for the two remaining
alternatives, until an alternative meets the upper threshold P (marked “Final selection”). At this
time, the alternative whose qtx (A) hits the upper threshold is selected as best, and sampling stops.
3.2. The BIZ Procedure with Heterogeneous and Unknown Sampling Variances
While Section 3.1 assumed a common known sampling variance
2
x
=
2
, sampling variances are
often heterogeneous and unknown in practice. In this section, we generalize the BIZ procedure to
handle heterogeneous sampling variances, in both variance-known and variance-unknown settings.
When the variances are known and are integer multiples of a common value, this procedure
retains the IZ guarantee of the known common-variance BIZ procedure. The continuous-time version of this procedure presented in Section 4.5 also retains the IZ guarantee, with a tight worst-case
preference-zone PCS bound. However, in discrete time, when the variances are unknown or lack
an integer multiple structure, we do not have a proof that it satisfies the IZ guarantee. Instead, in
this setting, we present this procedure as a heuristic.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
10
The discrete-time BIZ procedure for unknown and/or heterogeneous sampling variances is given
below in Alg. 2. Rather than taking only one sample from each alternative in contention for each
increment in t, Alg. 2 takes a variable number, storing the number of samples taken as ntx . We
let Ztx = Yntx ,x be the sum of all of these samples. The algorithm also maintains an adaptive
estimate b2tx of the sampling variance for alternative x, and is designed to keep ntx approximately
proportional to b2tx .
Alg. 2 accepts additional parameters beyond those accepted by Alg. 1: an integer n0 , and a
collection of integers B1 , . . . , Bk . n0 is the number of samples to use in a first stage of samples, for
which we recommend the value of 100. The parameter Bx governs the number of samples taken
from alternative x in each stage. In practice we recommend setting Bx to 1, and we leave it as a
free parameter because doing so supports theoretical analysis in Section 4.5.
Alg. 2 can also be used when the sampling variances are known. In this case, we set n0 = 0
and replace the estimators b2tx in Alg. 2 with their known values. If the variances are known and
identical, and we set B1 , . . . , Bk to their recommended values of 1, we recover Alg. 1.
Alg. 2 uses the quantity qbtx (A), defined here in terms of another quantity
qbtx (A) = exp
✓
Ztx
t
ntx
◆
X
exp
x0 2A
✓
Ztx0
t
ntx0
◆
,
t
P
t.
ntx0
0
= Px 2A
.
b2
x0 2A
(3)
tx0
Motivation We motivate Alg. 2 as follows. Consider what would happen if ntx were exactly
proportional to b2tx and our estimate b2tx =
2
x
were perfect, so ntx =
2
xt
consider the stochastic process Y 0 = (Ytx0 : t = 0, 1, 2, . . .), where Ytx0 = Ztx /
for some
2
x.
> 0. Then
A straightforward
computation shows that this stochastic process is a random walk whose increments are normal
with mean µx and variance 1/ . This variance does not depend on x, so to find arg maxx µx , we
may use a common-variance R&S procedure, such as Alg. 1.
Alg. 2 is derived by running Alg. 1 on Y 0 if these idealized conditions are met, or on an approximation to it if they are not. This approximation is Ytx0 ⇡ Ztx /( ntx / t ). Applying (2), but with
this approximation of Ytx0 in place of Ytx and the variance 1/ of the increments of Ytx0 in place of
2
, provides (3). When, as in the motivating situation described above, b2tx =
2
tx
and ntx =
2
tx t
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
11
Algorithm 2 Discrete-time implementation of BIZ, for unknown and/or heterogeneous variances.
Require: c 2 [0, 1
(P ⇤ ) k
1
1
],
> 0, P ⇤ 2 (1/k, 1), n0
integers. Recommended choices are c = 1
sampling variances
1:
2
x
(P ⇤ ) k
0 an integer, B1 , . . . , Bk strictly positive
1
1
, B1 = · · · = Bk = 1 and n0 = 100. If the
are known, replace the estimators b2tx with the true values
n0 = 0. To compute qbtx (A), use (3).
For each x, sample alternative x n0 times and set n0x
P ⇤, t
{1, . . . , k }, P
Let A
3:
while maxx2A qbtx (A) < P do
4:
5:
6:
7:
8:
9:
10:
11:
and set
n0 . Let Z0x and b20x be the sample
mean and sample variance respectively of these samples. Let t
2:
2
x,
0.
1.
while minx2A qbtx (A)  c do
Let x 2 arg minx2A qbtx (A).
Let P
P/(1
qbtx (A)).
Remove x from A.
end while
Let z 2 arg minx2A ntx /b2tx .
⇣
⌘
For each x 2 A, let nt+1,x = ceil b2tx (ntz + Bz )/b2tz .
For each x 2 A, if nt+1,x > ntx , take nt+1,x
ntx additional samples from alternative x. Let
Zt+1,x and b2t+1,x be the sample mean and sample variance respectively of all samples from
alternative x thus far.
12:
Increment t.
13:
end while
14:
Select x̂ 2 arg maxx2A Ztx /ntx as our estimate of the best.
for some
> 0, we have
t
= t = ntx /
2
x,
terms cancel, and Ytx0 is exactly equal to its approx-
imation. Below, in Section 4.5, we analyze special cases in which this occurs and show that, in
these situations, Alg. 2 satisfies the IZ guarantee, and does so with a tight bound on worst-case
preference-zone PCS in continuous time.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
12
3.3. BIZ Recovers PB⇤ as a Special Case
When c = 0 and variances are known and common, the discrete-time BIZ procedure is equivalent
to the non-elimination procedure PB⇤ introduced by Bechhofer et al. (1968). This can be seen as
follows: Because Ytx is almost surely finite at any fixed t, qtx (A0 ) > 0 almost surely. Thus, in Step 3
of Alg. 1, minx2A qtx (A) > 0 = c, the loop from Steps 4 to 7 will never execute, and A = {1, . . . , k }
and P = P ⇤ . The resulting procedure is then no longer an elimination procedure, and takes one
sample from every alternative at each point in time. It stops and selects the alternative with the
largest sample mean at the first time t for which maxx=1,...,k qtx ({1, . . . , k })
P ⇤ . This is exactly
the PB⇤ procedure of Bechhofer et al. (1968).
The parameter c determines the trade-o↵ between the number of stages of sampling and the
overall number of samples taken. When c = 0 the procedure does no elimination. When c > 0, the
procedure eliminates alternatives, with larger c causing more aggressive elimination, decreasing
the number of samples and increasing the number of stages. In highly parallel computing environments, and some biological and agricultural applications, one can evaluate many alternatives
simultaneously, and the focus is on minimizing the number of stages. In simulation, however, when
the number of alternatives is large compared to the parallelism in one’s computing environment,
the focus is on minimizing the number of samples taken.
In Section 5 we compare the number of samples taken by BIZ with c at its maximum value,
1
(P ⇤ ) k
c=1
1
1
, to the number taken by PB⇤ (which is BIZ with c at its minimum value, 0). Setting
(P ⇤ ) k
1
1
dramatically reduces the expected number of samples taken, particularly when
some alternatives are much worse than others. For use in simulation in a non-parallel setting, we
recommend c = 1
(P ⇤ ) k
1
1
.
4. Theoretical Analysis
In this section, we present our theoretical results: that IZ guarantees hold for Alg. 1 and, if variances
are known and have a special structure, Alg. 2; and that continuous-time generalizations of these
procedures have tight bounds on worst-case preference-zone PCS.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
13
First, we present a generalization of the R&S problem that allows both continuous-time and
discrete-time sampling in Section 4.1. Then, we consider the setting with common known variance:
Section 4.2 generalizes Alg. 1 to the continuous-time setting; Section 4.3 presents preliminary
definitions and results; and Section 4.4 presents the main theoretical results for the common known
variance setting. Section 4.5 considers the setting with heterogeneous known variance, presenting
first a continuous-time analogue of Alg. 2, and then theoretical results for both this continuous-time
analogue and Alg. 2 itself.
4.1. Generalization to Continuous Time: Observation Process
Although the R&S problem occurs in discrete time in practice, our theoretical analysis relies on
a generalization in which observations occur in continuous time, while decisions about eliminating
alternatives or stopping to select the best are made at any of a predetermined set of decision points.
When this set of decision points is the non-negative integers Z+ = {0, 1, . . .}, we recover the discretetime BIZ procedure. When it is the non-negative reals R+ = [0, 1), we obtain a continuous-time
version of BIZ, which we later show has tight worst-case bounds on preference-zone PCS.
Recall from Section 2 that, in discrete time and under Pµ, , the sum of all samples from alternative
x by time t, given in the stochastic process (Ytx : t 2 Z+ ), is a discrete-time random walk with
N (µx ,
2
x)
increments. We generalize this by letting (Ytx : t 2 R+ ) be a Brownian motion under Pµ,
starting from 0, with drift µx , volatility
x,
and independence across x. This is consistent with
the previous definition of Ytx at integer times t, since (Ytx : t 2 Z+ ) continues to be a discrete-time
random walk with N (µx ,
2
x)
increments. As before, for each A ✓ {1, . . . , k } we let YtA = (Ytx : x 2
A), and Yt = (Yt1 , . . . , Ytk ). We let F = (Ft : t 2 R+ ) be the filtration generated by (Yt : t 2 R+ ).
In the continuous-time setting, we assume that the variances are known. If they are not known,
they can be estimated with perfect accuracy from a sample path (Ytx : 0  t  ✏) for any ✏ > 0.
4.2. Generalization to Continuous Time: The BIZ Procedure with Common Variance
We now generalize the BIZ procedure for known common variance in discrete time (Alg. 1) to
include the continuous-time setting. In this section, we assume
2
x
=
2
for all x, with
2
known.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
14
This generalized BIZ procedure includes a parameter T which is a set of decision points, and is set
to either R+ or Z+ . Setting T = R+ provides a continuous-time procedure, while setting T = Z+
recovers the discrete-time procedure Alg. 1.
To define this generalized BIZ procedure, we recursively define a sequence of stopping times
0 = ⌧ 0  ⌧ 1  · · ·  ⌧k
1
 1, random variables Z1 , . . . , Zk
1
and P0 , P1 , . . . , Pk 1 , and random sets
A0 , A1 , . . . , Ak 1 . We first define ⌧0 , A0 and P0 as
P0 = P ⇤ ,
⌧0 = 0,
Then, for each n = 0, 1, . . . , k
A0 = {1, . . . , k }.
(4a)
2, we define ⌧n+1 , Zn+1 , An+1 and Pn+1 recursively given ⌧n , An ,
and Pn as
⇢
⌧n+1 = inf t 2 T \ [⌧n , 1) : min qtx (An )  c or max qtx (An )
x2An
x2An
Zn+1 2 arg min q⌧n+1 ,x (An ),
x2An
Pn ,
(4b)
An+1 = An \ {Zn+1 },
✓
◆
Pn+1 = Pn
1 min q⌧n+1 ,x (An ) .
x2An
Finally, with these quantities defined, x̂ is the single alternative in Ak 1 ,
x̂ 2 Ak 1 .
In this definition, the times ⌧1 , . . . , ⌧k
variables Z1 , . . . , Zk
1
1
(4c)
are times at which alternatives are eliminated, the random
are the alternatives eliminated at these times, and An is the set of alternatives
in contention starting at time ⌧n . The random variables P0 , . . . , Pk
1
are thresholds our posterior
probability of being best must achieve to allow us to stop sampling. The times at which alternatives
are eliminated has a particular structure: Initially, elimination occurs because minx2An qtx (An )  c,
i.e., because this posterior probability of being best fell below the lower threshold c. Eventually
though, an elimination occurs because maxx2An qtx (An )
Pn , i.e., because an alternative’s posterior
probability of being best exceeded the upper threshold Pn . At this time, Lemma 6 below shows
that all alternatives except one are eliminated simultaneously, ⌧n+1 = ⌧n+2 = · · · = ⌧k 1 , and that
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
the one alternative remaining in Ak
qtx (An )
1
15
(which is selected as best) is the one whose qtx (An ) satisfied
Pn at time t = ⌧n+1 . We define a random variable M so that the time at which this occurs
is ⌧M = ⌧n+1 , and the M th through the (k
1)st eliminations occur simultaneously in this way.
When T = Z+ , the algorithm defined by (4) is identical to Alg. 1. We see this as follows. The
stopping time ⌧n is the value of t in Alg. 1 at the time when the nth alternative is eliminated, either
explicitly in Step 6 (if n < M ), or implicitly in Step 10 (if n
M ). For 1  n < M , at time ⌧n , we
go through the inner while loop (Steps 2 through 7) for the nth time. The alternative x chosen for
elimination in Step 4 is Zn , Step 5 takes P from Pn
1
to Pn , and Step 6 takes A from An
1
to An .
At time t = ⌧M , the condition of the outer while loop checked in Step 2 fails to be satisfied, and
Alg. 1 goes to Step 10 and selects arg maxx2A Ytx = arg maxx2A Y⌧M ,x , which Lemma 6 below shows
is the same as the selection x̂ made by the procedure defined by (4).
In (4b), the choice among the arg min set for Zn+1 does not a↵ect the analysis, but we set Zn+1
to be the alternative with the smallest index in that set. When ⌧n+1 = 1, we choose Pn+1 = 1 and
Zn+1 uniformly at random from An , although again this choice does not a↵ect the analysis because
later in Lemma 4 we show that ⌧n+1 < 1 almost surely under any Pµ, . The claim that each ⌧n is
a stopping time is justified in Section 4.3, in Lemma 5.
This definition of BIZ in continuous time also provides an extension of PB⇤ to continuous time,
obtained by setting c = 0. The resulting procedure can be simplified, as is shown below in Lemma 7,
to a procedure that samples from all alternatives, eliminating none, until a selection is made at
time ⌧ = ⌧1 = ⌧2 = · · · = ⌧k
⌧ = inf
1
(
of x̂ 2 arg maxx=1,...,k Y⌧,x . This time can be written
)
k
X
2
⇤
2
t 2 T : max exp Ytx /
P
exp Ytx0 /
.
x
x0 =1
When T = Z+ , this is the original discrete-time PB⇤ procedure of Bechhofer et al. (1968).
4.3. Preliminaries for the Proofs
In this section, we present definitions and preliminary results that support our main theoretical
results for the common variance setting in Section 4.4. We continue to assume that
x, with
2
known.
2
x
=
2
for all
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
16
We first construct a probability measure Q under which the vector of sampling means is chosen
at random according to a prior distribution. To indicate that the vector of sampling means under
Q is random, and may di↵er from the true vector of sampling means µ, we denote it by ✓. We
emphasize that Q is a mathematical construct that we use to analyze the BIZ procedure, and is
di↵erent from the true sampling distribution Pµ, . We construct Q as follows. Let X ⇤ be chosen
uniformly at random from among 1, . . . , k, and let ✓X ⇤ = . Let ✓x = 0 for all x 6= X ⇤ . Configurations
of the form ✓X ⇤
✓x =
for some parameter
literature, and slippage configurations in which
> 0 are called slippage configurations in the R&S
=
are often the most difficult configurations
under which to select correctly.
We then define a family of probability measures that includes and generalizes Q. For each u 2 Rk
with u[k] 6= u[k
1] ,
define a probability measure Qu as follows. First, let (R(1), . . . , R(k)) under
Qu be a uniformly distributed permutation of (1, . . . , k). Then, let ✓x = uR(x) almost surely under
Qu , and let X ⇤ 2 arg maxx ✓x . (This argmax is unique because u[k] 6= u[k
1] .)
Given ✓, we let each
(Ytx : t 2 R+ ) be an independent Brownian motion under Qu with drift ✓x and volatility . Defining
u = [ , 0, . . . , 0], we have Q = Qu , so this definition generalizes the previously defined Q.
The following lemma provides an expression for the posterior probability that a specified alternative x0 has the largest sampling mean, given the prior Qu and partial information about the
permutation R. Proofs of this and other lemmas may be found in the appendix.
Lemma 1. Suppose
2
x
=
2
> 0 8x. Let A ✓ {1, . . . , k } with x0 2 A. Let u 2 PZ( ) and r⇤ 2
arg maxx ux . Let R be the (random) set of permutations r with R(x) = r(x) for x 2
/ A. Then,
⇤
0
Qu {X = x | YtA , (R(x))x2A
/ }=
X
r2R:r(x0 )=r⇤
exp
1 X
2
x2A
Ytx ur(x)
!
X
r2R
exp
1 X
2
!
Ytx ur(x) .
x2A
If R contains no permutations r with r(x0 ) = r⇤ , then the numerator in this expression is 0.
Lemma 1 has as a consequence Lemma 2 below, which gives the posterior probability under
Q that alternative x has the best sampling mean, given that the best is in a specified set A.
This expression is exactly qtx (A) (with x0 in place of x), defined earlier in (2). As discussed in
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
17
Section 3.1, this interpretation of qtx (A) motivates the BIZ procedure, and as we will see later,
plays an important role in its analysis.
Lemma 2. Suppose
2
x
=
2
⇤
> 0 8x. Let A ✓ {1, . . . , k } and x0 2 A. Then,
0
⇤
Q {X = x | YtA , X 2 A} = exp
✓
2
Ytx0
◆
X
exp
x2A
✓
2
◆
Ytx .
The expression is una↵ected if we also condition on (R(x))x2A
/ .
Later, we also use the following monotonicity result. Its proof involves algebraic manipulations
of expressions from Lemmas 1 and 2.
Lemma 3. Suppose
2
x
=
2
> 0 8x. Fix u 2 PZ( ), a permutation r0 of the integers {1, . . . , k }, and
y 2 Rk . Let A be a non-empty subset of {1, . . . , k }. Let B denote the event that X ⇤ 2 A, Ytx = yx
for each x 2 A, and R(x) = r0 (x) for each x 2
/ A. Then
max Qu {X ⇤ = x | B }
max Q {X ⇤ = x | B } ,
(5)
min Qu {X ⇤ = x | B }  min Q {X ⇤ = x | B } .
(6)
x2A
x2A
x2A
x2A
Our results also require the following pair of technical lemmas. The first states that, with probability 1, the BIZ procedure takes finitely many samples. Its proof employs a standard geometric
decay argument. The second states that the elimination times are stopping times of the filtration
generated by the observation process Y , and uses elementary manipulations of events.
Lemma 4. Suppose
2
x
=
2
> 0 8x. Then ⌧n < 1 a.s. under Pµ, for n = 0, 1, . . . , k
Lemma 5. For each n = 0, 1, . . . , k
1 and µ 2 Rk .
1, ⌧n defined by (4) is a stopping time of F .
In Section 4.2, we stated that, at the first elimination time ⌧n+1 caused by an alternative’s qtx (An )
exceeding Pn , all other alternatives are eliminated simultaneously, and this alternative is selected
as the best. We describe this behavior more formally in the following lemma, whose statement uses
the definition of the random variable M ,
⇢
M = inf n = 1, . . . , k
1 : max q⌧n ,x (An 1 )
x2An 1
Pn
1
,
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
18
so that ⌧M is the first time at which we eliminate an alternative because maxx2An
1
q⌧n ,x (An 1 )
exceeds Pn 1 .
2
x
Lemma 6. Suppose
2
=
> 0 8x. Then, for any µ 2 Rk , the following statements hold almost
surely under Pµ, .
(a) M  k
1.
M and x̂ 2 arg maxx2AM
(b) ⌧n = ⌧M for all n
(c) If T = R+ , ⌧M
1
1
Y⌧M ,x .
< ⌧M .
Lemma 6 allows us to formally state the previously claimed simplification of BIZ when c = 0,
from the discussion in Section 4.2 of the PB⇤ procedure.
Lemma 7. When c = 0, we have ⌧1 = ⌧2 = · · · = ⌧k
⌧ = inf
(
t 2 T : max exp
x
Ytx /
1
= ⌧ and x̂ 2 arg maxx=1,...,k Y⌧,x , where
2
P
⇤
k
X
exp
Ytx0 /
x0 =1
2
)
.
We may now state two lemmas, which together constitute the proof of the main result in Section 4.4. These lemmas use CS = {x̂ 2 arg maxx ✓x , ⌧k
1
< 1} to denote the event of correct selec-
tion. Lemma 8 shows that the non-Bayesian probability of correct selection PCS(u, ) is identical
to the probability of correct selection under the Bayesian prior Qu . The proof follows a symmetry
or “equalizing” argument.
Lemma 8. Suppose
2
x
2
=
> 0 8x. Let u 2 Rk . Then Qu {CS} = PCS(u, ). Furthermore,
PCS(u, ) is invariant to translations and permutations of u.
Lemma 9 shows that the conditional Bayesian PCS is bounded below by the random variable
Pn 1 , with equality in the case of continuous-time sampling and prior Q.
Lemma 9. Suppose
2
x
=
2
> 0 8x. Then, for each n = 1, . . . , k
⇤
Qu CS | F⌧n , (R(x))x2A
/ n 1 , X 2 An 1 , M
1 and each u 2 PZ( ),
n
If T = R+ and u = [ , 0, . . . , 0] then this inequality holds with equality.
Pn 1 .
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
19
4.4. Theoretical Results for the Common Variance Setting
Using the preliminary results from the previous section, we now state and prove our main result
common known variances, Theorem 1. The first statement shows that the BIZ procedure satisfies
the IZ guarantee in both discrete and continuous time. The second statement in the theorem shows
that, in continuous time, the bound on the worst-case preference-zone PCS of the BIZ procedure
is tight. This second statement can be interpreted as showing that any slack in the PCS bound for
BIZ in discrete time is due to the gap between the time at which the continuous-time procedure
would eliminate an alternative and the next integer-valued time.
Theorem 1. Let c 2 [0, 1
(P ⇤ ) k
1
1
], T 2 {Z+ , R+ },
> 0, P ⇤ 2 (1/k, 1), and
2
x
=
2
> 0 for all
x. Then, under the BIZ procedure defined by (4),
P⇤
PCS(µ, )
for all µ 2 PZ( ).
Furthermore, if T = R+ ,
inf
µ2PZ( )
Proof:
PCS(µ, ) = P ⇤ .
Let µ 2 PZ( ). Let µ0 be a permutation of µ such that µ01
given by ux = µ0x
µ01 + . Thus, u1 =
µ0x for all x. Let u 2 Rk be
and ux  0 for all x 6= 1. Because u is a permutation and
translation of µ, Lemma 8 implies
PCS(µ, ) = Qu {CS}.
Furthermore, u 2 PZ( ). Since X ⇤ 2 A0 and M
(7)
1 with probability 1, and the complement of A0
is empty, taking n = 1 in Lemma 9 shows
Qu {CS | F⌧1 } = Qu {CS | F⌧1 , X ⇤ 2 A0 , M
1}
P0 = P ⇤ .
(8)
Then, the tower property of conditional expectation provides
Qu {CS} = EQu [Qu {CS | F⌧1 }]
EQu [P ⇤ ] = P ⇤ ,
where EQu is the expectation under Qu . Combining (7) and (9) provides PCS(µ, )
(9)
P ⇤.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
20
We have shown that PCS(µ, )
P ⇤ for all µ 2 PZ( ). This shows that inf µ2PZ( ) PCS(µ, )
P ⇤.
To see that the infimum is equal to P ⇤ when T = R+ , consider µ = u = [ , 0, . . . , 0]. Lemma 9 shows
that, in this case, the inequalities in (8) and (9) are actually equalities. Combining equality in
(9) with the equality (7) shows that PCS([ , 0, . . . , 0], ) = P ⇤ , implying inf µ2PZ( ) PCS(µ, )  P ⇤ .
This shows that the infimum must in fact equal P ⇤ .
⇤
The last paragraph of the proof shows that the infimum of the PCS over the preference zone is
achieved by the configuration [ , 0, . . . , 0]. The invariance of the PCS to translations and permutations of the configuration shown by Lemma 8 implies that the infimum is also attained by any
slippage configuration with parameter . Thus, these slippage configurations are least-favorable for
the BIZ procedure with common known variance.
4.5. Theoretical Results for the Heterogeneous Variance Setting
We now discuss the heterogeneous known variance setting. We present a continuous-time procedure that is analogous to the discrete-time Alg. 2. This continuous-time procedure satisfies the IZ
guarantee and has tight worst-case preference-zone PCS bounds. We then use this fact to show
that Alg. 2 satisfies the IZ guarantee when variances are known and have a special integer multiple
structure.
For each x let nx (t) =
2
x t.
This quantity is the continuous-time analogue of the discrete-
time quantity ntx in Alg. 2, and in the certain special cases discussed below, nx (t) = ntx for all
integer times t. Now define a stochastic process (Ytx0 : t
0) as Ytx0 = Ynx (t),x /
2
x.
A straightforward
computation shows that (Ytx0 : t 2 R+ ) is a Brownian motion with drift µx and volatility 1/ , so
any algorithm that performs R&S in the continuous-time common-variance case can be run on
the modified observation processes Y 0 , and the result is a continuous-time R&S algorithm for the
original observation process Y . This was also noted for discrete time in Section 3.2.
With this motivation, the continuous-time BIZ procedure for known heterogeneous variances is
obtained by applying the continuous-time BIZ procedure for common sampling variances from (4)
to the modified observation process Y 0 . More explicitly, this procedure is defined by first setting
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
P0 = P ⇤ ,
⌧0 = 0,
21
A0 = {1, . . . , k },
(10a)
then defining recursively, for n = 0, 1, . . . , k 2,
⇢
0
0
⌧n+1 = inf t 2 T \ [⌧n , 1) : min qtx
(An )  c or max qtx
(An )
x2An
x2An
Pn ,
Zn+1 2 arg min q⌧0 n+1 ,x (An ),
x2An
(10b)
An+1 = An \ {Zn+1 },
✓
◆
0
Pn+1 = Pn
1 min q⌧n+1 ,x (An ) .
x2An
0
where qt,x
(A) is obtained by substituting Y 0 in place of Y and 1/ in place of
0
qt,x
(A)
= exp (
Ytx0 )
X
exp (
Ytx0 0 )
= exp
x0 2A
✓
2
x
Ynx (t),x
◆
X
exp
x0 2A
✓
2
x0
2
Y
,
n0x (t),x0
◆
,
(10c)
and finally letting the selected alternative x̂ be the single entry in Ak 1 ,
x̂ 2 Ak 1 .
When sampling variances are identically equal to
(10d)
2
across alternatives and
= 1/
2
, we have
0
nx (t) = t, Ytx0 = Yt,x , qt,x
(A) = qt,x (A), and the procedure defined by (10) is identical to (4).
The following theorem shows that this procedure satisfies the IZ guarantee, and its worst-case
preference-zone PCS bound is tight in continuous time. This is true even when sampling variances
di↵er from each other.
Theorem 2. Let c 2 [0, 1
(P ⇤ ) k
1
1
], T = {Z+ , R+ },
> 0, P ⇤ 2 (1/k, 1), and
2
1
> 0, . . . ,
2
k
> 0.
Then, under the BIZ procedure defined by (10),
PCS(µ, )
P⇤
for all µ 2 PZ( ).
Furthermore, if T = R+ ,
inf PCS(µ, ) = P ⇤ .
µ2PZ( )
Proof:
and
Let x̂ be the selection decision x̂ defined by (10), with the specified values for P ⇤ , ,c,T,
2
1, . . . ,
2
k.
We use the superscript
to emphasize that x̂ assumes sampling variances
2
x.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
22
Let x̂ be the selection decision defined by (4) using
2
= 1/ , and the same specified values of
P ⇤ , , c, and T. The distribution of Y 0 under Pµ, is equal to the distribution of Y under Pµ, .
Consequently, the distribution of x̂ under Pµ, is equal to the distribution of x̂ under Pµ, .
This implies that Pµ,
x̂ 2 arg maxx µx = Pµ, {x̂ 2 arg maxx µx }. The result then follows from
applying Theorem 1 to Pµ, {x̂ 2 arg maxx µx }.
⇤
While (10) is directly implementable in continuous time, it is more difficult to apply in discrete
time. While one can set T to Z+ in (10), the resulting procedure is not always implementable in
discrete time. The reason is that (10) requires observations of Ynx (t),x for t 2 T. If nx (t) =
2
xt
can
fail to be an integer for some t 2 T, then these observations may be unavailable in discrete time.
However, if the variances have a special integer multiple structure, then (10) is implementable
in discrete time, and is equivalent to Alg. 2. In particular, suppose the variances
satisfy
2
x
= ax
2
for some common
2
and integers a1 , a2 , . . . , ak . If we set
2
x
= 1/
are known and
2
and T = Z+ ,
then nx (t) = ax t is always an integer for t 2 Z+ , and all observations of Ynx (t),x required by (10)
are available in discrete time. Furthermore, in this case, (10) is identical to Alg. 2 with parameters
Bx = ax , n0 = 0, and b2x =
2
x.
A direct consequence of this and Theorem 2 is that Alg. 2 satisfies
the IZ guarantee, in this special case. We have just shown the following corollary to Theorem 2.
Corollary 1. Let
2
x
= ax
2
for all x, where ax 2 Z+ with ax
1,
2
> 0. Let c 2 [0, 1
(P ⇤ ) k
1
1
],
> 0, and P ⇤ 2 (1/k, 1). Then, under the BIZ procedure for known heterogeneous sampling variances given in Alg. 2 with n0 = 0, Bx = ax , and b2tx =
PCS(µ, )
P⇤
2
x
for all x,
for all µ 2 PZ( ).
Outside of the common variance setting, the integer multiple structure assumed by Corollary 1
is unlikely to appear in practice. Also, in practice one would set Bx to 1, rather than to the values
assumed by Corollary 1, to improve the responsiveness of the algorithm and reduce expected sample
sizes. Thus, while Corollary 1 provides insight into the behavior of Alg. 2, it is not designed to
provide an IZ guarantee that directly applies to how this algorithm is used in practice. Instead, we
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
23
present Alg. 2 as a heuristic in practical settings, and we use numerical experiments to investigate
its statistical properties in the next section.
5. Numerical Results
We demonstrate the performance of the BIZ procedure in discrete time with maximum elimination
(c = 1
(P ⇤ ) k
1
1
) on standard test problems, and compare it to another leading IZ procedure,
the KN procedure of Kim and Nelson (2001), first on problems with common known sampling
variance, then on problems with common unknown sampling variance, and finally on problems
with heterogeneous unknown sampling variance.
The KN procedure improves over previously proposed IZ procedures in a number of configurations (Kim and Nelson 2001), and the KN family of procedures has been regarded by Kim and
Nelson (2006) and Branke et al. (2007) as state-of-the-art for IZ R&S. Improvements of the original
KN procedure, particularly the KN++ procedure of Goldsman et al. (2002) and the KVP and
UVP procedures of Hong (2006), o↵er better performance than KN in some settings with unknown
and/or heterogeneous sampling variance, but KN remains one of the best existing IZ procedures.
In problems with known variance, we modify KN from its original version in Kim and Nelson
(2001) to take advantage of knowing the variance. Where the original procedure uses estimates of
the variance, the modified procedure uses the actual value. The modified procedure also uses the
tighter constant (h⇤ )2 = 2c⌘ ⇤ in place of the parameter h2 from Kim and Nelson (2001), where c = 1,
and ⌘ ⇤ satisfies g(⌘ ⇤ ) = 12 exp( ⌘ ⇤ ) =
1 P⇤
.
k 1
We set n0 = 1. This modified procedure is the same as
the P procedure in Wang and Kim (2012), and when used in common variance configurations, the
same as Paulson’s procedure Paulson (1964). In problems where the sampling variance is unknown,
we use KN as originally described in Kim and Nelson (2001), with c = 1 and n0 = 100.
We also compare to the PB⇤ procedure of Bechhofer et al. (1968), which is BIZ with no elimination,
as described in Section 3.3. In our figures, we denote the PB⇤ procedure by BKS, the initials of the
authors of Bechhofer et al. (1968).
We examine both the PCS and the expected total number of samples taken, denoted E[N ]. We
emphasize that N counts the total number of samples taken, and so a procedure without elimination
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
24
has E[N ] = kE[⌧k 1 ], while a procedure with elimination has E[N ]  kE[⌧k 1 ]. Rather than plotting
E[N ] directly, we plot the expected number of samples taken divided by the number of alternatives,
E[N ]/k. This normalizes E[N ] and clarifies performance trends.
Figure 2 shows the performance of KN, BKS, and BIZ (Algorithm 1), under three di↵erent
configurations with common known variance, described in more detail below. Each row shows
performance under a di↵erent configuration vs. the number of alternatives k. Left-hand panels
show E[N ]/k and right-hand panels show PCS, obtained using 10,000 independent replications.
SC: Row 1 of Figure 2 shows performance under a slippage configuration (SC), in which µ1 =
and µx = 0 for x 6= 1. For many procedures, including BIZ and KN, a slippage configuration with
parameter
is the configuration in the preference zone in which correctly selecting the best is most
difficult, and is often used as a test case to better understand the behavior of R&S procedures.
Here,
= 1,
x
=
2
= 100, and P ⇤ = 0.9. Experiments were performed at k = 2, 3, . . . , 8, and then
at integral powers of 2 up to k = 214 = 16, 384.
PCS under this SC is shown in the right-hand panel of Row 1. We know from their IZ guarantees
that PCS for all three procedures is bounded below by the target probability P ⇤ = 0.9. Apparent
deviations below 0.9 are due to estimation error — standard errors for PCS reported for BIZ and
p
BKS are approximately 0.9 ⇥ 0.1/104 = .003. Under KN, which has a loose worst-case preferencezone PCS bound and which over-delivers on PCS for large problems, PCS quickly rises away from
P ⇤ as the number of alternatives grows. In contrast, under BIZ and BKS, PCS remains close
to P ⇤ . The proximity of PCS to the target shows that, although the lower bound on worst-case
preference-zone PCS given in Theorem 1 is no longer tight as we move from continuous to discrete
time, the bound remains nearly tight in discrete time, at least in the settings tested.
E[N ]/k under this SC is shown in the left-hand panel of Row 1. Points plotted have standard error
less than 2 for KN and BIZ, and less than 7 for BKS. As the number of alternatives grows large,
KN begins taking a very large number of samples, while the number of samples taken by BIZ grows
at a much slower rate. For the largest problem considered, k = 16, 384, KN takes (22.4 ± 0.2) ⇥ 106
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
1
KN
BKS
1500 BIZ
2000
0.98
0.96
PCS
E[N]/k
25
1000
KN
BKS
BIZ
0.94
0.92
500
0.9
0
0.88
1
10
100
1000
10000
1
300
1
250
0.98
200
100
0.94
0.9
0
0.88
1
10
100
1000
10000
1
Number of alternatives (k)
2200
2000
1800
1600
1400
1200
1000
800
600
400
200
0
10
100
1000
10000
Number of alternatives (k)
1
KN
BIZ
0.95
PCS
E[N]/k
10000
0.92
50
0.9
KN
BIZ
0.85
0.8
0.75
1
10
100
1000
10000
1
Number of alternatives (k)
Figure 2
1000
KN
BKS
BIZ
0.96
KN
BKS
BIZ
150
100
Number of alternatives (k)
PCS
E[N]/k
Number of alternatives (k)
10
10
100
1000
10000
Number of alternatives (k)
Common known variance: Rows shows PCS and E[N ]/k vs. the number of alternatives k under the
following configurations: SC (row 1); MDM (row 2); and RPI (row 3). For SC and MDM, P ⇤ = 0.9. For
RPI, P ⇤ = 0.8. Sampling variances are common and known.
samples in expectation, while BIZ takes (7.574 ± .003) ⇥ 106 . The number of samples taken by KN
is 3 times larger than the number taken by BIZ.
Although both BIZ and BKS deliver a PCS that is close to the target, BIZ requires many fewer
samples than BKS (a factor of 4.4 times fewer samples at k = 16, 384) because of its ability to
eliminate poor alternatives early. This ability and its associated improvement in sampling efficiency,
while clearly present, is relatively modest here in this slippage configuration where all of the
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
26
suboptimal alternatives have the same true sampling mean. In the configuration to be examined
next, the ability to eliminate alternatives plays a much larger role.
MDM: Row 2 of Figure 2 shows performance under the monotone decreasing means (MDM)
configuration, in which µx =
x. Unlike a slippage configuration in which the best is exactly
better than the second best, where it is difficult to tell the best alternative from the others, the
MDM configuration allows easier identification of the best alternative. As with the SC considered
previously, = 1,
x
=
2
= 100, P ⇤ = 0.9, and experiments were performed at k = 1, 2, 3, . . . , 8, and
then at integral powers of 2 up to k = 214 = 16, 384.
In the right-hand panel of Row 2, we again see that KN’s PCS grows quickly away from the target,
and is indistinguishable from 1 for k > 100 on the scale at which the figure is plotted. In contrast,
BKS’s PCS stays close to the target. BIZ’s PCS is further from the target for intermediate values
of k than it was under the SC, which shows that BIZ over-delivers on PCS when the configuration
is not least favorable. However, BIZ’s over-delivery is not extreme, as its PCS stays below 0.99
over all k, and comes back to the target for large values of k.
The left-hand panel of Row 2 shows E[N ]/k in the MDM configuration. Points plotted have
standard error less than 2 in all cases (shrinking to less than 0.02 for k
512 for BIZ and KN). Here,
as in the SC, BIZ outperforms KN across all values of k, although by a smaller margin. At k=16,384,
KN requires 34, 242 ± 17 samples in expectation, 1.5 times more than the estimated 22,170 required
by BIZ (the standard error on BIZ’s estimated sample size is less than .001 ⇥ 16, 384 = 16.4). BIZ
also performs significantly better than BKS across all values of k. These aspects of the behavior
under the MDM configuration are similar to those seen under the SC.
The left-hand panel of Row 2 also shows a number of di↵erences between MDM and SC. Because
the alternatives that are added to the MDM configuration as k increases are progressively further
from the best, an MDM configuration with 10, 000 alternatives is not much more difficult than one
with 100 alternatives. Thus, as the number of alternatives grows large in the MDM configuration,
the number of samples E[N ] taken by a good procedure should grow very slowly. This is the reason
why the total number of samples required per alternative, E[N ]/k, which initially rises under all
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
27
three procedures, eventually decreases under BIZ and KN. In contrast, the average number of
samples required per alternative under BKS reaches a plateau and does not drop, causing the
sampling e↵ort to grow linearly in k. This is because BKS is unable to eliminate bad alternatives,
and so must sample all of them until the stopping time ⌧ . Beyond approximately 10 alternatives,
adding more alternatives does not increase E[⌧ ] under BKS, but it does significantly increase E[N ]
because E[⌧ ] = kE[N ]. This demonstrates the importance of elimination.
RPI: Row 3 of Figure 2 shows performance of KN and BIZ on random problem instances (RPI).
To generate a single random problem instance, we first choose k = ceil(exp(U )), where U is uniform
between 0 and log(16, 384), and ceil(x) is the smallest integer greater than or equal to x. When
generated in this way, k
2 with probability 1. We then generated the sampling mean of each of
the k alternatives randomly from an independent normal distribution with mean 0 and variance
0.25. Those sampling means that are not best, but are within
to be exactly
2
= 100,
of the best, are then rounded down
from the best. This ensures that the configuration is in the preference zone. Here,
= 0.5, and P ⇤ = 0.8. We generated 75 random problem configurations in this way.
The right-hand panel of Row 3 shows that PCS under both procedures for each problem instance
is at or above P ⇤ = 0.8, as predicted by the theory. (There is one configuration with an estimated
BIZ PCS of 0.797 ± .004, which is below 0.8, but not by a statistically significant margin.) In this
RPI setting (as in the SC and MDM settings), KN’s PCS is larger than that of the BIZ procedure,
and it becomes extremely close to 1 as k increases. On several large problem configurations, KN
selected correctly on every one of the 10, 000 independent replications. This indicates extreme
over-delivery. In contrast, BIZ’s PCS is lower than that of KN, and is evenly distributed within
the interval [P ⇤ , 1] = [0.8, 1] for small k. As k increases, BIZ’s typical PCS moves toward 1, with
typical estimated PCS for the largest problems near 0.995. All estimated BIZ PCS values are less
than 0.997. While closer to 1 than under SC and MDM, BIZ’s over-delivery on PCS under these
random problem instances is still less than that of KN.
The left-hand panel of Row 3 shows E[N ]/k vs. k for both BIZ and KN. Standard errors are all
less than 5. BIZ requires many fewer samples than KN, especially for large problems. As k grows,
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
28
KN-UNK
BIZ-UNK
2000
PCS
E[N]/k
1500
1000
500
0
1
10
100
1000
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
10000
KN-UNK
BIZ-UNK
1
Number of alternatives (k)
300
PCS
E[N]/k
200
150
100
50
0
1
10
100
1000
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
10000
1
100
1000
10000
0.95
1500
PCS
E[N]/k
10
1
KN-UNK
2000 BIZ-UNK
1000
500
0.9
0.85
0.8
0
KN-UNK
BIZ-UNK
0.75
10
10000
Number of alternatives (k)
2500
1
1000
KN-UNK
BIZ-UNK
Number of alternatives (k)
100
1000
10000
1
Number of alternatives (k)
Figure 3
100
Number of alternatives (k)
KN-UNK
BIZ-UNK
250
10
10
100
1000
10000
Number of alternatives (k)
Common unknown variance: Rows shows PCS and E[N ]/k vs. the number of alternatives k under the
following configurations: SC (row 1); MDM (row 2); and RPI (row 3). For SC and MDM, P ⇤ = 0.9. For
RPI, P ⇤ = 0.8. Sampling variances are common and unknown.
E[N ]/k under BIZ initially increases with k, taking an average value near 800 at k = 10, and then
declines for k
10 down to near 600 for the largest problems. In contrast, the average number of
samples taken by KN is near 1400 at k = 10, and rises to near 2000 at k = 16, 384. Here again,
BIZ’s ability to deliver a PCS that is closer to the target allows it to take fewer samples.
Common unknown variance: Figure 3 shows the performance of KN and BIZ (Algorithm 2)
on the configurations SC, MDM, and RPI, with unknown sampling variance. Although the vari-
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
29
ance is common across the alternatives, this fact is unknown to the algorithms. In this situation,
we set n0 = 100 for both KN and BIZ, and set Bx = 1 for BIZ. Quantities are estimated using
10, 000 independent replications. Performance trends from the common known variance setting are
apparent here as well: compared with KN, BIZ has a PCS that is closer to the target of P ⇤ = 0.9,
and takes fewer samples. One di↵erence appears in the MDM configuration: the limiting value of
E[N ]/k as k grows is n0 = 100, because each alternative must be sampled at least n0 times. These
experimental results show that BIZ works well, even when the variance is unknown, at least in the
settings investigated.
Heterogeneous unknown variance: Figure 4 shows the performance of KN and BIZ (Algorithm 2
on problems with heterogeneous and unknown variance. The problem configurations considered
are the following modifications of previously considered configurations. First, the Slippage Configuration with Increasing Variance (SC-INC) uses the same sampling means as the SC considered
previously, µ1 =
2
x
= (1 +
x 1
) 2,
k 1
and µx = 0 for x 6= 1, and heterogeneous increasing sampling variances, with
so that
2
k
=2
DEC) is analogous, but uses
2
x
2
1.
The Slippage Configuration with Decreasing Variance (SC-
= (1 +
k x
) 2,
k 1
so
2
1
=2
2
k.
Second, the Monotone Decreasing
Means Configuration with Increasing Variance (MDM-INC) uses the same means as MDM, and the
same variances as SC-INC. Similarly, the Monotone Decreasing Means Configuration with Decreasing Variance (MDM-DEC) uses the same means as MDM, and the same variances as SC-DEC.
Third, Random Problem Instances with Heterogeneous Variances (RPI-HET) uses the same set of
randomly generated means as RPI, and the same sampling variances as SC-INC and MDM-INC.
In each configuration, performance quantities were estimated with 10, 000 independent replications. Again, BIZ performs well on these problem configurations, providing a PCS closer to P ⇤
and taking fewer samples than KN, providing evidence that BIZ’s performance is robust to both
unknown and heterogeneous variance.
Additional numerical experiments: Additional numerical experiments investigating the probability of good selection (Nelson and Banerjee 2001) for configurations outside the preference zone
are presented in the appendix.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
30
2500
1500
PCS
E[N]/k
KN-UNK
2000 BIZ-UNK
1000
500
0
1
10
100
1000
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
10000
KN-UNK
BIZ-UNK
1
Number of alternatives (k)
2500
1500
PCS
E[N]/k
KN-UNK
2000 BIZ-UNK
1000
500
0
1
10
100
1000
10000
PCS
E[N]/k
200
100
0
10
100
1000
10000
PCS
E[N]/k
200
100
0
10
100
1000
1
1000
10000
10
100
1000
10000
1
0.95
2000
PCS
E[N]/k
100
Number of alternatives (k)
KN-UNK
3000 BIZ-UNK
2500
1500
1000
0.9
0.85
0.8
500
0
KN-UNK
BIZ-UNK
0.75
10
10000
KN-UNK
BIZ-UNK
Number of alternatives (k)
100
1000
Number of alternatives (k)
Figure 4
10
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
10000
3500
1
1000
Number of alternatives (k)
300
1
100
KN-UNK
BIZ-UNK
1
KN-UNK
BIZ-UNK
400
10
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
Number of alternatives (k)
500
10000
Number of alternatives (k)
300
1
1000
KN-UNK
BIZ-UNK
1
KN-UNK
BIZ-UNK
400
100
1
0.98
0.96
0.94
0.92
0.9
0.88
0.86
Number of alternatives (k)
500
10
Number of alternatives (k)
10000
1
10
100
1000
10000
Number of alternatives (k)
Heterogeneous unknown variance: Row shows PCS and E[N ]/k as a function of the number of alternatives k under the following configurations: SC-INC (row 1); SC-DEC (row 2); MDM-INC (row 3);
MDM-DEC (row 4); and RPI-HET (row 5). For SC-INC, SC-DEC, MDM-INC, and MDM-DEC, P ⇤ =
0.9. For RPI-HET, P ⇤ = 0.8. Sampling variances are heterogeneous and unknown.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
31
6. Conclusion
We have developed a new IZ procedure, called the Bayes-inspired IZ (BIZ) procedure. In continuous
time, our lower bound on its worst-case preference-zone probability of correct selection is tight,
and in discrete time, numerical experiments demonstrate that our lower bound is close to the
procedure’s true worst-case probability of correct selection. This is the first sequential elimination
procedure with tight worst-case preference-zone bounds on probability of correct selection for more
than 2 alternatives.
These theoretical results assume that the sampling variances are known, and have a particular
integer multiple structure. In practice, sampling variances are unknown and do not possess an
integer multiple structure, but numerical experiments suggest that the procedure’s performance is
robust to violations of these assumptions.
The tightness of this lower bound allows the procedure to take fewer samples than other IZ
procedures, especially for problems with large numbers of alternatives. While the BIZ procedure
takes only as many samples as is required to meet the desired lower bound on probability of
correct selection, other procedures like the KN procedure of Kim and Nelson (2001) deliver a true
probability of correct selection that is much larger than requested, and consequently take many
more samples than is needed. Thus, having a tight lower bound improves efficiency and allows the
BIZ procedure to select the best with fewer samples.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
32
Appendix A: Proofs
The suppositions ur⇤ = maxx ux and u 2 PZ( ) imply ur⇤ > ux for all x 6= r⇤ .
Proof of Lemma 1.
This implies that the event X ⇤ = x0 is identical to the event R(x0 ) = r⇤ .
Before computing the probability that R(x0 ) = r⇤ , we first compute the probability that R =
r for a generic r. Consider any fixed permutation r of the integers {1, . . . , k }. By Bayes rule,
the conditional probability Qu {R = r | YtA , (R(x))x2A
/ } is proportional to the likelihood of YtA and
(R(x))x2A
given R = r. If r 2
/ R, this likelihood is 0 because r is inconsistent with the observed
/
(R(x))x2A
/ . If r 2 R, then r is consistent with the observed (R(x))x2A
/ , and this likelihood is the
same as the likelihood of YtA given R = r. This likelihood is
Qu {YtA 2 dy | R = r } = (2⇡
2
)
1 |A|
2
exp
"
2
1 X
= c(r, y) exp
2
1 X
2
)
1 |A|
2
exp
h
1
2 2
P
fixed values for (r(x))x2A
/ , it follows that
yx
ur(x)
x2A
yx ur(x)
x2A
where c(r, y) = (2⇡
2
!
2
#
dy
dy,
i
2
2
y
+
u
x
r(x) . Because r 2 R is a permutation of u with
x2A
P
2
x2A ur(x)
is identical for all r 2 R. Furthermore,
does not depend upon r, so c(r, y) is identical for all r 2 R.
P
2
x2A yx
This allows us to rewrite the likelihood as
1 X
Qu {YtA 2 dy | R=r} / exp
2
!
yx ur(x) dy.
x2A
Now let r vary over R. Because the event {R(x0 ) = r⇤ } is the union of all the events {R = r}
with r(x0 ) = r⇤ , and only those r 2 R have nonzero likelihood, we have
Qu {R(x) = 1 | YtA , (R(x))x2A
/ }=
X
r2R:r(x0 )=r⇤
exp
1 X
2
x2A
Ytx ur(x)
!
X
r2R
exp
1 X
2
x2A
!
Ytx ur(x) ,
which recovers the claimed expression.
Proof of Lemma 2.
Recall that Q = Qu , where u = [ , 0 . . . , 0] 2 PZ( ). Let R be as defined in
Lemma 1 and let r⇤ = 1 2 arg maxx ux . Then, the expression from Lemma 1 provides an expression
for Qu {X ⇤ = x0 | YtA , (R(x))x2A
/ }.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
On the event X ⇤ 2 A, and for r 2 R,
the alternative r
1
P
x2A Ytx ur(x)
33
= Yt,r
1 (r )
⇤
(r⇤ ) which is best under permutation r, and ur
since ur(x) = 0 for all x except
= . The event X ⇤ 2 A is
1 (r )
⇤
⇤
measurable given (R(x))x2A
/ A. On this event, the expression
/ because X 2 A i↵ R(x) 6= r⇤ for all x 2
from Lemma 1 becomes
⇤
Qu {X ⇤ = x0 | YtA , (R(x))x2A
/ , X 2 A} =
X
exp
r2R:r(x0 )=r⇤
X
=
exp
r2R:r(x0 )=r⇤
= ax0 exp
✓
2
✓
✓
Ytx0
Y 0
2 tx
2
◆
Ytx0
◆
X
r2R
◆
X
exp
X
✓
Y
2 t,r
X
1 (r )
⇤
exp
x2A r2R:r(x)=r⇤
ax exp
x2A
✓
2
◆
✓
◆
2
Ytx
◆
Ytx ,
where ax = | {r 2 R : r(x) = r⇤ } | is the number of elements of R in which a given alternative x is
1)! is constant for all x 2 A on the event X ⇤ 2 A, we cancel it to obtain
the best. Since ax = (|A|
⇤
0
⇤
Qu {X = x | YtA , (R(x))x2A
/ , X 2 A} = exp
✓
2
Ytx0
◆
X
exp
x2A
✓
2
◆
Ytx .
Finally, this expression does not depend upon (R(x))x2A
/ , and so
⇤
0
⇤
Qu {X = x | YtA , X 2 A} = exp
Proof of Lemma 3.
✓
2
Ytx0
◆
X
x2A
exp
✓
2
◆
Ytx .
Without loss of generality we assume that maxx ux = . If this is not the
case, then we add the constant
maxx0 ux0 to each ux , which does not change the value of
Q u {X ⇤ = x | B }.
Let x̃ 2 arg maxx0 2A yx0 . By the expression in Lemma 2, maxx2A Q {X ⇤ = x | B } = Q {X ⇤ = x̃ | B }.
We will show that Q {X ⇤ = x̃ | B }  Qu {X ⇤ = x̃ | B }, which is sufficient to show (5) because
Qu {X ⇤ = x̃ | B }  maxx2A Qu {X ⇤ = x | B }.
For x 2 A, define fu (x) = Qu {X ⇤ = x | B } /Qu {X ⇤ = x̃ | B }. This quantity is well-defined because
Qu {X ⇤ = x̃ | B } > 0. Then,
Qu {X ⇤ = x̃ | B } = P
Qu {X ⇤ = x̃ | B }
1
=P
.
⇤
x2A Qu {X = x | B }
x2A fu (x)
P
Similarly define f (x) = Q {X ⇤ = x̃ | B } /Q {X ⇤ = x | B } so that Q {X ⇤ = x̃ | B } = [ x2A f (x)] 1 .
P
Since [ x2A zx ] 1 is decreasing in each zx , to show that Qu {X ⇤ = x̃ | B } Q {X ⇤ = x̃ | B }, it is
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
34
enough to show that fu (x)  f (x) for each x 2 A. This follows trivially for x = x̃, so we now consider
x 6= x̃.
Let x0 2 A with x0 6= x̃. We have, by Lemmas 2 and 1, where R and r⇤ are as defined in Lemma 1,
P
P
1
yx ur(x)
2
1
r2R:r(x̃)=r⇤ exp
Px2A
=P
1
0
f (x )
r2R:r(x0 )=r⇤ exp( 2
x2A yx ur(x) )
1
fu (x0 )
exp
exp(
yx̃
2 yx 0 )
2
Multiplying through by the strictly positive quantity
2
4
X
1 X
exp
2
r2R:r(x0 )=r⇤
x2A
shows that the sign of this expression for [fu (x0 )]
X
exp
2
yx 0 +
r2R:r(x̃)=r⇤
=
X
X
j6=r⇤ r2R:r(x̃)=r⇤ ,r(x0 )=j
=
X
2
1 X
2
yx ur(x)
x2A
0
4exp @
2
0
exp @
X
j6=r⇤ r2R:r(x̃)=r⇤ ,r(x0 )=j
0
exp @
!
yx 0 +
!3
✓
yx ur(x) 5 exp
[f (x0 )]
1
X
exp
1
2
is the same as the sign of
yx̃ +
r2R:r(x0 )=r⇤
u r⇤
2
yx̃ +
uj
1 X
X
1
2
x2A\{x̃,x0 }
2
yx̃ +
u r⇤
2
yx 0 +
uj
2
2
yx̃ +
1
2
X
y +
2 x̃
X
1
2
x2A\{x̃,x0 }
1
1
!
yx ur(x) A
x2A\{x̃,x0 }
y 0+
2 x
yx ur(x)
x2A
yx 0 +
2
yx 0
2
◆
13
yx ur(x) A5
h
⇣u
⌘
j
yx ur(x) A exp 2 yx0
exp
⇣u
j
y
2 x̃
⌘i
.
(11)
The second line uses that the values of r(x) for x 2 A \ {x̃, x0 } are the same between the two sums
P
r2R:r(x̃)=r⇤ ,r(x0 )=j
and
P
r2R:r(x0 )=r⇤ ,r(x̃)=j .
The third line uses that ur⇤ = .
Consider the sign of the last term, exp(uj yx0 /
u 2 PZ( ) imply uj  0. Furthermore, yx̃
implies exp(yx0 uj /
2
)
exp(Ytx̃ uj /
2
)
2
)
exp(uj yx̃ /
2
). Together ur⇤ = , j 6= r⇤ and
yx0 then implies exp(uj yx̃ /
2
)  exp(uj yx0 /
2
), which
0. Thus, the sign of (11) is nonnegative, which shows that
fu (x)  f (x) for each x, which shows (5).
The proof of (6) follows a similar argument, but with x̃ 2 arg minx2A yx . In this case, fu (x)
for each x 2 A because yx̃  yx0 implies that exp(yx0 uj /
positive.
2
)
exp(yx̃ uj /
2
f (x)
)  0 and (11) is non-
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
Proof of Lemma 4.
35
We will show recursively that ⌧n < 1 a.s. for each n = 0, 1, . . . , k
statement holds for ⌧0 = 0. Now suppose that ⌧n < 1 a.s. for some n  k
1. The
2, and we will show
that ⌧n+1 < 1 a.s.
2
Define a random variable a =
1
log((k
Pn )). If n = 0 then Pn = P ⇤ 2 (0, 1),
1)Pn )/(1
implying a is finite. If n > 0, then P ⇤  Pn  P ⇤ /(1
c)n  P ⇤ /(1
c)k
2
 P ⇤ /(P ⇤ )(k
2)/(k 1)
< 1,
implying Pn 2 (0, 1) and a is finite.
Fix a deterministic time t
0, let ⌧ 0 = ⌧n + t, and let x̃ be any F⌧ 0 -measurable random variable
that is almost surely in An . Consider the event that ⌧n < 1 and Y⌧ 0 +1,x̃
x 2 An \ {x̃}. On this event, q⌧ 0 +1,x (An )/q⌧ 0 +1,x̃ (An ) = exp(
x 2 An \ {x̃} and
2
q⌧ 0 +1,x̃ (An ) = 41 +
X
x2An \{x̃}
3
2
Y⌧ 0 +1,x
a for each
Y⌧ 0 +1,x̃ ))  exp( a /
(Y⌧ 0 +1,x
2
) for
1
q⌧ 0 +1,x (An )/q⌧ 0 +1,x̃ (An )5
⇥
1 + (k
1) exp( a /
2
)
⇤
1
= Pn .
The value of a was chosen to make this last equality with Pn hold. Thus, on the event considered,
q⌧ 0 +1,x̃
Pn and ⌧n+1  ⌧ 0 + 1.
We now define x̃ 2 arg maxx2An Y⌧ 0 ,x , which is F⌧ 0 -measurable and is almost surely in An . The
previous discussion implies
Pµ, {⌧n+1  ⌧ 0 + 1 | F⌧ 0 , ⌧n+1 > ⌧ 0 }
Pµ, {Y⌧ 0 +1,x̃
a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 } .
Y⌧ 0 +1,x
This implies that
Pµ, {⌧n+1  ⌧ 0 + 1 | F⌧ 0 , ⌧n+1 > ⌧ 0 }
Pµ, {Y⌧ 0 +1,x̃
Y⌧ 0 ,x̃ , Y⌧ 0 ,x̃
Y⌧ 0 +1,x
a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 }
Pµ, {Y⌧ 0 +1,x̃
Y⌧ 0 ,x̃ , Y⌧ 0 ,x
Y⌧ 0 +1,x
= Pµ, {Y⌧ 0 +1,x̃
Y⌧ 0 ,x̃ | F⌧ 0 }
Y
a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 }
where the second line follows from the fact Y⌧ 0 ,x̃
x2An \{x̃}
Pµ, {Y⌧ 0 ,x
Y⌧ 0 +1,x
a | F⌧ 0 } .
Y⌧ 0 ,x and the third line follows from the inde-
pendence of increments of Ysx under Pµ, .
The probability Pµ, {Y⌧ 0 +1,x̃
variable, Y⌧ 0 +1,x̃
Y⌧ 0 ,x̃ | F⌧ 0 } is the probability of a conditionally N (µx̃ ,
Y⌧ 0 ,x̃ , exceeding 0. This probability is
(µx̃ /
2
2
) random
), which is bounded below by
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
36
(minx µx /
Pµ, {Y⌧ 0 ,x
Yt
2
Y⌧ 0 +1,x
Yt,x
1,x
maxx0 µx0 )/
). Here,
is the normal cumulative distribution function. Similarly, the probability
a | F⌧ 0 } is the probability of a conditionally N ( µx
a, exceeding 0. This probability is ( (a+µx )/
2
2
a,
2
) random variable,
), and is bounded below by ( (a+
).
Thus, replacing ⌧ 0 with ⌧n + t,
Pµ, {⌧n+1  ⌧n + t + 1 | F⌧n +t , ⌧n+1 > ⌧n + t}
2
(min µx /
x
)
h
( (a + max µx )/
x
2
)
ik
1
.
Let ✏ be the quantity on the right-hand side of this inequality. ✏ < 1 is an F⌧n -measurable random
variable that does not depend on t.
By repeated application of this inequality, we have for any finite t that Pµ, {⌧n+1 > ⌧n + t | F⌧n } 
(1
✏)t , which vanishes in the limit as t ! 1. This implies that Pµ, {⌧n+1 = 1} = 0.
Proof of Lemma 5.
We show the claim using induction. It is trivially true for ⌧0 = 0. Now
suppose the claim is true for some n, and we show that it holds for n + 1. To show that ⌧n+1 is a
stopping time, it suffices to show that {⌧n+1  t} is Ft measurable for each t
0. Let t
0, and
write
{⌧n+1  t} = {⌧n  t} \ {qsx (An ) 2
/ (c, Pn ) for some s 2 [⌧n , t] \ T }
=
[
A✓{1,...,k}
{⌧n  t} \ {An = A} \ {qsx (A) 2
/ (c, Pn ) for some s 2 [⌧n , t] \ T }
The last event in this expression {qsx (A) 2
/ (c, Pn ) for some s 2 [⌧n , t] \ T } can be rewritten as
S
s2Q {s
2 [⌧n , t] \ T} \ {qsx (A) 2
/ (c, Pn )} where Q denotes the rational numbers. For the case T =
Z+ , this follows from T ⇢ Q. For the case T = R+ , it follows from the fact that (c, Pn ) is an open set,
and from the continuity of the sample paths of s 7! qsx (A). (This sample-path continuity follows in
turn from the continuity of qsx (A) considered as a function of Ys , and the continuity of the sample
paths of Ys .) It is convenient to rewrite and summarize this as
{qsx (A) 2
/ (c, Pn ) for some s 2 [⌧n , t] \ T } =
[
s2Q\T\[0,t]
{⌧n  s} \ {qsx (A) 2
/ (c, Pn )}
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
37
Now consider the term {qsx (A) 2
/ (c, Pn )}. This can be rewritten as
{qsx (A) 2
/ (c, Pn )} =
\
{p
p2Q
Pn } [ {qsx (A) 2
/ (c, p)} .
Putting this all together,
{⌧n+1  t} =
[
\
A✓{1,...,k} p2Q
s2Q\T\[0,t]
[
=
\
A✓{1,...,k} p2Q
s2Q\T\[0,t]
{⌧n  t} \ {An = A} \ {⌧n  s} \ {p
Pn } [ {qsx (A) 2
/ (c, p)} ,
{⌧n  t, An = A} \ {⌧n  s} \ {⌧n  t, p
Pn } [ {qsx (A) 2
/ (c, p)} .
In this expression, {⌧n  t, An = A} is Ft -measurable because An is F⌧n measurable; {⌧n  t, p
Pn }
is Ft -measurable because Pn is F⌧n measurable; {⌧n  s} is Ft -measurable because ⌧n is a stopping
time and s  t; and {qsx (A) 2
/ (c, p)} is Ft -measurable because qsx (A) is Fs -measurable and s  t.
Since countable unions and intersections of Ft -measurable events are Ft -measurable, {⌧n+1  t}
is Ft -measurable.
Part (a): Consider the event {M
Proof of Lemma 6.
M =k
k
1}. On this event, we show that
Pk 2 .
(12)
1. To show this, it is sufficient to show
max q⌧k
1 ,x
x2Ak 2
We know ⌧k
1
(Ak 2 )
< 1 almost surely by Lemma 4, and the definition of ⌧k
1
implies that either (12)
holds or the following relation holds,
min q⌧k
x2Ak 2
The number of elements in Ak
2
1 ,x
(Ak 2 )  c.
(13)
is exactly 2. Call them x(1) and x(2) , where Y⌧k
1 ,x
(1)
Y⌧k
1 ,x
(2)
.
By Lemma 2, for x 2 x(1) , x(2) ,
q⌧ k
and q⌧k
1 ,x
(1)
q⌧ k
1 ,x
(2)
1 ,x
(Ak 2 ) =
exp(
exp(
2
Y⌧k
Y⌧k 1 ,x )
(1) ) + exp( 2 Y⌧
1 ,x
k
2
1 ,x
(2)
)
◆
.
,
.
The expression (12) holds i↵
exp
✓
2
Y⌧k
1 ,x
(1)
◆
Pk 2
1 Pk
exp
2
✓
2
Y⌧k
1 ,x
(2)
(14)
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
38
Similarly, the expression (13) holds i↵
✓
◆
exp
Y
(1)
2 ⌧k 1 ,x
On the event {M
Pk
c
c
P⇤
(1 c)k
2
= (1
c)
P⇤
(1 c)k
1
where the first inequality is due to Pn = Pn 1 /(1
1, . . . , k
2 on {M
k
exp
✓
Y
2 ⌧k
1 ,x
(2)
◆
.
(15)
1},
k
2
1
 (1
c)
P⇤
(P ⇤ )(k
minx2An
1
1)/(k 1)
2

c,
q⌧n,x (An 1 ))  Pn /(1
1}, and the second inequality is due to c  1
Pk 2
1 Pk
=1
(P ⇤ )1/(k
1)
c) for n =
. This implies
1 c
1 c
=
,
1 (1 c)
c
implying that if (15) holds, so does (14). The fact that at least one of (15) and (14) holds implies
that (14) holds, which implies that (12) holds. This shows part (a).
Part (b): Let Y ⇤ = maxx2AM
1
Y⌧M ,x . We claim that, for n = M, M + 1, . . . , k
1, the following
statements hold:
(i) ⌧n = ⌧M .
(ii) maxx2An Y⌧n ,x = Y ⇤ .
(iii) maxx2An
1
q⌧n ,x (An 1 )
Pn 1 .
We show this by induction on n.
We first consider the base case, n = M . In this case, (i) follows trivially. (ii) follows because
AM = AM
1
\ {ZM }, ZM 2 arg minx2AM
1
Y⌧M ,x and |AM
1|
=k
(M
1)
2 by part (a). (iii)
follows from the definition of M .
We now show the induction step. Suppose that (i), (ii), and (iii) hold for some n 2 [M, k
2].
We show they also hold for n + 1. We have, by (ii) for n,
" P
# 1
exp(
Y
)
exp( 2 Y ⇤ )
exp( 2 Y ⇤ )
⌧
,x
2
n
max q⌧n ,x (An ) = P
=P
· P x2An
. (16)
x2An
exp(
Y
)
exp(
Y
)
exp(
Y
⌧
,x
⌧
,x
⌧n ,x )
2
2
2
n
n
x2An
x2An 1
x2An 1
P
The first term of (16) satisfies exp( 2 Y ⇤ )/ x2An 1 exp( 2 Y⌧n ,x ) = maxx2An 1 q⌧n ,x (An 1 )
Pn 1 , where we have used (iii) for n. We rewrite the ratio in the second term of (16) as
P
exp( 2 Y⌧n ,x )
exp( 2 Y⌧n ,Zn )
P x2An
=1 P
= 1 q⌧n ,Zn (An 1 ) = 1
min q⌧n ,x (An 1 ),
x2An 1
x2An 1 exp( 2 Y⌧n ,x )
x2An 1 exp( 2 Y⌧n ,x )
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
where we use that An
1
39
\ An = {Zn } in the first equality, the definition of qt,x (·) in the second
equality, and that Zn 2 arg minx2An
q⌧n ,x (An 1 ) in the third equality. Combining these two, (16)
1
becomes
max q⌧n ,x (An )
Pn 1 [1
x2An
1
min q⌧n ,x (An 1 )]
= Pn .
x2An 1
This and the definition of ⌧n+1 implies that ⌧n+1 = ⌧n , implying in turn that (i) and (iii) are met
for n + 1.
Finally, Zn+1 2 arg minx2An Y⌧n+1 ,x = arg minx2An Y⌧M ,x , and An+1 = An \ {Zn+1 }, so
maxx2An+1 Y⌧n+1 ,x = maxx2An+1 Y⌧M ,x = maxx2An Y⌧M ,x = Y ⇤ and (iii) is met for n + 1.
This shows that (i),(ii), and (iii) are met for n = M, . . . , k
in Ak 1 , and (i),(ii) for n = k
maxx2AM
1
1 implies that this element must have Y⌧k
Y⌧M ,x , so x̂ 2 arg maxx2AM
Part (c): To show that ⌧M > ⌧M
PM
1.
1. Finally, x̂ is the single element
= Y⌧M ,x̂ = Y ⇤ =
1 ,x̂
Y⌧M ,x . This shows part (b).
1
1,
it is sufficient to show that maxx2AM
1
q⌧ M
1 ,x
arg maxx2AM
Y⌧M
1 ,x
max q⌧M
1 ,x
2
x2AM 2
=P
exp(
x2AM
= q⌧ M
P
<
1 ,x̃
(AM
2)
Y⌧M 1 ,x̃ )
exp( 2 Y⌧M
2
1)
1
1
= AM
1
exp( 2 Y⌧M 1 ,x )
x2AM
2
exp( 2 Y⌧M 1 ,x )
q⌧ M
1 ,ZM
1
1 ,x
q⌧ M
x2AM
used that 1
= q⌧ M
2
(AM
=1
(AM
2)
)
1 ,x̃
=P
1 ,ZM
2
(AM
1
2)
exp(
Y⌧M 1 ,x̃ )
exp( 2 Y⌧M
1
x2AM
(AM
2)
\ {Z M
1}
exp
2 Y ⌧M
P
⇣
x2AM
=1
2
2
= q⌧M
1 ,x̃
1 ,x
)
(AM
P
·P
x2AM 1
x2AM 2
1 )PM
exp(
2
Y⌧M
1 ,x
)
exp(
2
Y⌧M
1 ,x
)
2 /PM
1.
in going from the second to the third line to write
1 ,ZM
1
⌘
exp( 2 Y⌧M 1 ,x )
minx2AM
2
q⌧M
1 ,x
=1
q⌧ M
(AM
2)
1 ,ZM
= PM
1
(AM
2 /PM
1
2 ).
1 ,x̃
(AM
1)
=
maxx2AM 2 q⌧M 1 ,x (AM
PM 2 /PM 1
where the inequality follows from the definition of M .
2)
<
PM
PM 2
2 /PM
= PM
1
We have also
to write the last
equality on the third line. From this we have,
q⌧ M
2. Let x̃ 2
, so
where we have used AM
that
1)
We show this by considering two cases. In the first case, suppose M = 1. Then the
result follows because maxx2A0 q0x = 1/k < P ⇤ = P0 . In the second case, suppose M
P
(AM
1,
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
40
Finally, because an alternative with minimal Y⌧M
we have arg maxx2AM
maxx2AM
1
q⌧M
1 ,x
(AM
2
Y⌧M
1)
1 ,x
= q⌧ M
= arg maxx2AM
1 ,x̃
(AM
1)
1
1 ,x
Y⌧M
< PM
1.
is eliminated in going from AM
1 ,x
2
to AM
1,
, and x̃ is an element of this set. Thus,
This shows part (c).
Proof of Lemma 7.
Because Ytx is almost surely finite at any fixed t, qtx (A0 ) > 0 almost surely. Thus, when c = 0,
minx2An qtx (An ) > 0 = c. This implies that ⌧1 = inf {t 2 T : maxx2A0 qtx (A0 )
P0 } and M = 1. By
Lemma 6, ⌧ = ⌧1 . Recalling P0 = P ⇤ , A0 = {1, . . . , k }, and the definition of qtx (A), allows us to
rewrite ⌧ as claimed.
Proof of Lemma 8.
The invariance of PCS(u, ) to permutations of u follows from the invari-
ance of the BIZ procedure to the labeling of the alternatives. Let a 2 R, and we now show that
PCS(u, ) = PCS(u + [a, . . . , a], ), so PCS(u, ) is also invariant to translations of u. We work
under the probability measure Pµ, with µ = u + [a, . . . , a]. Let Wtx = (Ytx
tux
ta))/ so that
{Wtx : t 2 R+ , x 2 {1, . . . , k }}, is a k-dimensional standard Brownian motion. We then write qtx (A)
in terms of Wtx . From (2), and Ytx = Wtx + tux + ta,
qtx (A) = exp
= exp
✓
✓
2
( Wtx + tux + ta)
( Wtx + tux )
2
◆
◆
X
x0 2A
X
x0 2A
exp
exp
✓
✓
2
( Wtx + tux + ta)
◆
(
W
+
tu
)
,
tx
x
2
where we have divided the numerator and denominator by exp( ta/
2
◆
). Thus, the path of the
stochastic process (qtx (A) : t 2 R+ , A ✓ {1, . . . , k }) can be written as a function of the path of a
k-dimensional standard Brownian motion, and this function does not depend upon a. Furthermore,
since the event of correct selection can be written entirely in terms of the paths of qtx (A), the
probability of correct selection PCS(u + [a, . . . , a], ) also does not depend upon a.
Now we show that Qu {CS} = PCS(u, ). First, rewrite the probability of CS under Qu as
Qu {CS} =
X
r
Qu {R = r} Qu {CS | R = r} .
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
41
Conditioning on R = r under Qu is the same as conditioning on ✓ = ũ where ũx = ur(x) for each
x. Thus, Qu {CS | R = r} = Qu {CS | ✓ = ũ} = Pũ, {CS } = PCS(ũ, ). Since PCS(·, ) is invariant to
permutations of the alternatives, and ũ is a permutation of u, PCS(ũ, ) = PCS(u, ). This implies
Qu {CS} =
X
r
Qu {R = r} PCS(u, ) =
X
r
!
Qu {R = r} PCS(u, ) = PCS(u, ),
because Qu {R = r} sums over r to 1.
Proof of Lemma 9.
Fix u 2 PZ( ). First, since ⌧k
1
< 1 almost surely under any Pv, with v 2
Rk by Lemma 4, and Qu is a mixture of such Pv, , Qu {CS} = Qu {x̂ = X ⇤ , ⌧k
1
< 1} = Qu {x̂ = X ⇤ }.
Thus we may work with the probability of {x̂ = X ⇤ } rather than the probability of CS.
To simplify notation, let Gn be the sigma-algebra generated by F⌧n and (R(x))x2A
/ n 1 . Let Cn be
the event {X ⇤ 2 An 1 , M = n} and Dn be the event {X ⇤ 2 An 1 , M > n}. Both events are measurable with respect to Gn . To show the lemma, it is enough to show
Qu {x̂ = X ⇤ | Gn , Cn }
Pn
1
for n = 1, . . . , k
1,
(17)
Qu {x̂ = X ⇤ | Gn , Dn }
Pn
1
for n = 1, . . . , k
2,
(18)
with equality if u = [ , 0, . . . , 0] and T = R+ . We first show (17), and then show (18).
Proof of (17). On the event Cn ✓ {M = n}, Lemma 6 implies x̂ 2 arg maxx2An
Lemma 1, this argmax set is identical to arg maxx2An
1
1
Y⌧n ,x . By
Qu {X ⇤ = x | Gn }. This implies that,
Qu {x̂ = X ⇤ | Gn , Cn } = max Qu {X ⇤ = x | Gn , Cn }
x2An 1
max Q {X ⇤ = x | Gn , Cn } = max q⌧n ,x (An 1 )
x2An 1
x2An 1
Pn 1 .
On the second line, the first inequality is due to Lemma 3, the equality is due to the definition of
qtx (A), and the second inequality is due to the definition of M and Cn ✓ {M = n}.
If u = [ , 0, . . . , 0], then Qu = Q and the first inequality is an equality. If T = R+ , the second
inequality is an equality due to the continuity of the paths of Brownian motion and ⌧M
1
< ⌧M , as
shown in Lemma 6. Thus, if u = [ , 0, . . . , 0] and T = R+ , then Qu {x̂ = X ⇤ | Gn , Cn } = Pn 1 .
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
42
Proof of (18). We now show (18). The proof is by backward induction on n. We first consider
n=k
2. Lemma 6 part (a) implies that M = k
(17) implies (18) holds when n = k
Now fix some n < k
1 almost surely on the event M > k
2. Then,
2, and with equality if u = [ , 0, . . . , 0] and T = R+ .
2. Our induction hypothesis is that (18) holds for n + 1, and that it holds
with equality if u = [ , 0, . . . , 0] and T = R+ . We will show this is then also true for n. We have,
Qu {x̂ = X ⇤ | Gn , Dn } = Qu {Zn 6= X ⇤ | Gn , Dn } Qu {x̂ = X ⇤ | Gn , Dn , Zn 6= X ⇤ } .
(19)
We rewrite the first term in (19) as
Qu {Zn 6= X ⇤ | Gn , Dn } = 1
=1
1
=1
Q u {Z n = X ⇤ | G n , D n }
min Qu {x = X ⇤ | Gn , Dn }
x2An 1
min Q {x = X ⇤ | Gn , Dn }
x2An 1
min q⌧n ,x (An 1 ) = Pn 1 /Pn .
x2An 1
The second line follows from Zn 2 arg minx2An
1
Qu {x = X ⇤ | Gn }. The third line follows from
Lemma 3, and is equality if u = [ , 0, . . . , 0]. The fourth and last line follows from the definition of
⌧n and that we condition on Dn ✓ {M > n}.
We now consider the second term of (19). Define the event
E = Dn \ {Zn 6= X ⇤ } = {X ⇤ 2
/ An 1 , Zn 6= X ⇤ , M > n} = {X ⇤ 2 An , M
n + 1} = Cn+1 [ Dn+1 .
The second term of (19), Qu {x̂ = X ⇤ | Gn , Dn , Zn 6= X ⇤ }, can be rewritten,
Qu {x̂ = X ⇤ | Gn , E } = EQu [Qu {x̂ = X ⇤ | Gn+1 , E } | Gn , E]
EQu [Pn | Gn , E] = Pn ,
where EQu indicates the expectation taken with respect to Qu . In the first equality we have
used the tower property of conditional expectation and Gn ⇢ Gn+1 . In the inequality, we
have used that (17) implies Qu {x̂ = X ⇤ | Gn+1 , Cn+1 }
Qu {x̂ = X ⇤ | Gn+1 , Dn+1 }
Pn and the induction hypothesis implies
Pn , which together imply Qu {x̂ = X ⇤ | Gn+1 , Cn+1 [ Dn+1 }
inequality is equality if T = R+ and u = [ , 0, . . . , 0].
Pn . This
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
43
Combining the two terms of (19), we obtain
Pn 1
Pn = Pn ,
Pn
Qu {x̂ = X ⇤ | Gn , Dn }
with equality if u = [ , 0, . . . , 0] and T = R+ . Thus, the induction statement holds.
Appendix B: Additional Numerical Results
In Figure 5 we test the performance of BIZ on problem configurations outside the preference
zone. Recall that a procedure satisfying the indi↵erence-zone guarantee requires that the lower
bound P ⇤ on PCS only be met for problem configurations in the preference zone PZ( ) =
µ 2 Rk : µ[k]
µ[k
1]
, where the best alternative is better than the second best alternative by
at least . For studying problem configurations outside the preference zone, Nelson and Banerjee
(2001) suggests that performance be measured by the probability of good selection (PGS),
PGS(µ, ) = Pµ,
n
x̂
max µx
x
o
, ⌧ <1 .
For configurations in the preference zone, PGS is equal to PCS. But for configurations outside the
preference zone, the two quantities di↵er.
A natural generalization of the indi↵erence-zone guarantee would be a lower bound on PGS over
all configurations, i.e., a probability of good selection guarantee of the form,
PGS(µ, )
P⇤
for all µ 2 Rk+ .
A procedure that satisfies this PGS guarantee also satisfies the IZ guarantee, but a procedure
satisfying the IZ guarantee does not necessarily satisfy the PGS guarantee.
The numerical results in Figure 5 suggest that the BIZ procedure may satisfy the PGS guarantee,
at least in the common known variance case. However, this has not been confirmed theoretically,
and doing so is left for future work.
Acknowledgments
The author was supported by AFOSR YIP FA9550-11-1-0083 and NSF CAREER CMMI-1254298. The
author would like to thank Rolf Waeber and Shane Henderson for helpful discussions, and the associate
editor and several anonymous referees for helpful comments.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
PGS
44
0.99
0.98
0.97
0.96
0.95
0.94
0.93
0.92
0.91
0.9
0.89
k=3
k=10
k=100
0
0.5
1
1.5
2
(mu(1) - mu(2)) / Delta
Figure 5
The probability of good selection (PGS) plotted as a function of (µ1
µ2 )/ for BIZ procedure with
common known-variance for problem configurations with k = 3, 10, and 100 alternatives. All configurations considered had µ = [ , µ2 , 0, . . . , 0],
= 0.5, P ⇤ = 0.9, and
= 10. When (µ1
have a slippage configuration with parameter . When this quantity is
µ2 )/ = 1, we
1, the configuration is in the
preference zone, and when it is < 1, it is in the indi↵erence zone.
References
Andradóttir, S., S.H. Kim. 2010. Fully sequential procedures for comparing constrained systems via simulation. Naval Research Logistics 57(5) 403–421.
Bechhofer, R.E. 1954. A single-sample multiple decision procedure for ranking means of normal populations
with known variances. The Annals of Mathematical Statistics 25(1) 16–39.
Bechhofer, R.E., J. Kiefer, M. Sobel. 1968. Sequential Identification and Ranking Procedures. University of
Chicago Press, Chicago.
Bechhofer, R.E., T.J. Santner, D.M. Goldsman. 1995. Design and Analysis of Experiments for Statistical
Selection, Screening and Multiple Comparisons. J.Wiley & Sons, New York.
Berger, James O. 1985. Statistical decision theory and Bayesian analysis. 2nd ed. Springer-Verlag, New
York.
Branke, J., S.E. Chick, C. Schmidt. 2007. Selecting a selection procedure. Management Science 53(12)
1916–1932.
Chick, S.E., J. Branke, C. Schmidt. 2010. Sequential sampling to myopically maximize the expected value
of information. INFORMS Journal on Computing 22(1) 71–80.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
45
Chick, S.E., K. Inoue. 2001. New two-stage and sequential procedures for selecting the best simulated system.
Operations Research 49(5) 732–743.
Dieker, AB, Seong-Hee Kim. 2012. Selecting the best by comparing simulated systems in a group of three
when variances are known and unequal. Proceedings of the 2012 Winter Simulation Conference. IEEE.
Fabian, V. 1974. Note on Anderson’s sequential procedures with triangular boundary. The Annals of
Statistics 2(1) 170–176.
Frazier, P., W. B. Powell, S. Dayanik. 2008. A knowledge gradient policy for sequential information collection.
SIAM Journal on Control and Optimization 47(5) 2410–2439.
Frazier, P., W. B. Powell, S. Dayanik. 2009. The knowledge gradient policy for correlated normal beliefs.
INFORMS Journal on Computing 21(4) 599–613.
Frazier, P.I., W.B. Powell. 2008. The knowledge-gradient stopping rule for ranking and selection. S.J.
Mason, R.R. Hill, L. Mönch, O. Rose, T. Je↵erson, J.W. Fowler, eds., Proceedings of the 2008 Winter
Simulation Conference. Institute of Electrical and Electronics Engineers, Inc., Piscataway, New Jersey,
305–312.
Goldsman, D., S.H. Kim, W.S. Marshall, B.L. Nelson. 2002. Ranking and selection for steady-state simulation: Procedures and perspectives. INFORMS Journal on Computing 14(1) 2–19.
Gupta, S.S., K.J. Miescke. 1996. Bayesian look ahead one-stage sampling allocations for selection of the best
population. Journal of Statistical Planning and Inference 54(2) 229–244.
Hartmann. 1991. An improvement on Paulson’s procedure for selecting the population with the largest mean
from k normal populations with a common unknown variance. Sequential Analysis 10(1-2) 1–16.
Hartmann, M. 1988. An improvement on Paulson’s sequential ranking procedure. Sequential Analysis 7(4)
363–372.
Hong, J. 2006. Fully sequential indi↵erence-zone selection procedures with variance-dependent sampling.
Naval Research Logistics 53(5) 464–476.
Kim, Seong-Hee, AB Dieker. 2011. Selecting the best by comparing simulated systems in a group of three.
Proceedings of the 2011 Simulation Conference. IEEE, 3987–3997.
Kim, Seong-Hee, Barry L. Nelson. 2001. A fully sequential procedure for indi↵erence-zone selection in
simulation. ACM Trans. Model. Comput. Simul. 11(3) 251–273.
Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection
Article submitted to Operations Research; manuscript no. OPRE-2011-09-480
46
Kim, S.H., B.L. Nelson. 2006. Selecting the best system. S.G. Henderson, B.L. Nelson, eds., Handbook in
Operations Research and Management Science: Simulation. Elsevier, Amsterdam, 501–534.
Kim, S.H., B.L. Nelson. 2007. Recent advances in ranking and selection. Proceedings of the 39th conference
on Winter simulation: 40 years! The best is yet to come. IEEE Press, Piscataway NJ, 162–172.
Malone, Gwendolyn J, Seong-hee Kim, David Goldsman, Demet Batur. 2005. Performance of Variance
Updating Ranking and Selection Procedures. Proceedings of the 2005 Winter Simulation Conference.
IEEE.
Nelson, B., J. Swann, D. Goldsman, W. Song. 2001. Simple procedures for selecting the best simulated
system when the number of alternatives is large. Operations Research 49(6) 950–963.
Nelson, B.L., S. Banerjee. 2001. Selecting a good system: Procedures and inference. IIE Transactions 33(3)
149–166.
Paulson, E. 1964. A sequential procedure for selecting the population with the largest mean from k normal
populations. The Annals of Mathematical Statistics 35(1) 174–180.
Paulson, E. 1994. Sequential procedures for selecting the best one of k Koopman-Darmois populations.
Sequential Analysis 13(3).
Rinott, Y. 1978. On two-stage selection procedures and related probability-inequalities. Communications in
Statistics-Theory and Methods 7(8) 799–811.
Swisher, J.R., S.H. Jacobson, E. Yücesan. 2003. Discrete-event simulation optimization using ranking, selection, and multiple comparison procedures: A survey. ACM Transactions on Modeling and Computer
Simulation 13(2) 134–154.
Wang, Huizhu, Seong-Hee Kim. 2012. On the conservativeness of fully sequential indi↵erence-zone procedures. Submitted for publication .