Submitted to Operations Research manuscript OPRE-2011-09-480 A Fully Sequential Elimination Procedure for Indi↵erence-Zone Ranking and Selection with Tight Bounds on Probability of Correct Selection Peter I. Frazier School of Operations Research and Information Engineering, Cornell University, Ithaca, NY 14853, [email protected], http://people.orie.cornell.edu/pfrazier/ We consider the indi↵erence-zone (IZ) formulation of the ranking and selection problem with independent normal samples. In this problem, we must use stochastic simulation to select the best among several noisy simulated systems, with a statistical guarantee on solution quality. Existing IZ procedures sample excessively in problems with many alternatives, in part because loose bounds on probability of correct selection lead them to deliver solution quality much higher than requested. Consequently, existing IZ procedures are seldom considered practical for problems with more than a few hundred alternatives. To overcome this, we present a new sequential elimination IZ procedure, called BIZ (Bayes-inspired Indi↵erence Zone), whose lower bound on worst-case probability of correct selection in the preference zone is tight in continuous time, and nearly tight in discrete time. To the author’s knowledge, this is the first sequential elimination procedure with tight bounds on worst-case preference-zone probability of correct selection for more than two alternatives. Theoretical results for the discrete-time case assume that variances are known and have an integer multiple structure, but the BIZ procedure itself can be used when these assumptions are not met. In numerical experiments, the sampling e↵ort used by BIZ is significantly smaller than that of another leading IZ procedure, the KN procedure of Kim and Nelson (2001), especially on the largest problems tested (214 = 16, 384 alternatives). 1. Introduction In the use of simulation, one commonly encounters the problem of selecting the best among several simulated systems, e.g., selecting the method for operating a supply chain with minimum average cost, or selecting the configuration of an assembly line with maximum throughput. The higher-level problem of deciding how many simulation samples to take from each system to best support this selection of the best is called the ranking and selection (R&S) problem. Doing well in R&S requires balancing the amount of time spent sampling against the quality of the ultimate selection. 1 Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 2 We consider the indi↵erence-zone (IZ) formulation of the R&S problem, in which we wish to correctly select the best alternative with a probability exceeding a user-specified target, whenever the best alternative is sufficiently separated from the others. A sampling procedure having this property is said to satisfy the IZ guarantee. This problem has a rich history, dating to the seminal work Bechhofer (1954), with early work summarized in the monograph Bechhofer et al. (1968). Research in the area has been active since that time (see, e.g., Paulson (1964), Fabian (1974), Rinott (1978), Hartmann (1988), Paulson (1994), Nelson et al. (2001), Goldsman et al. (2002), Hong (2006), Andradóttir and Kim (2010)). This large body of research is summarized in Bechhofer et al. (1995), and more recent work is reviewed in Swisher et al. (2003), Kim and Nelson (2006, 2007). The goal in designing IZ sampling procedures is to take as few samples as possible while still satisfying the IZ guarantee. While early IZ procedures introduced in Bechhofer (1954), Paulson (1964), Fabian (1974), Rinott (1978), Hartmann (1988, 1991), Paulson (1994) satisfy the IZ guarantee, they provide a probability of correct selection (PCS) much larger than the user-specified target probability. This is due in part to their use of the Bonferonni inequality, which leads to loose theoretical PCS bounds, and is problematic because it leads them to sample more than necessary. This issue becomes more severe as the number of alternatives grows. More recent procedures developed in Kim and Nelson (2001), Goldsman et al. (2002), Hong (2006) have better performance, but these procedures continue to use the Bonferonni inequality, causing them to be overly conservative in large problems, in the sense that they sample more than necessary and over-deliver on PCS targets (Branke et al. 2007). Recent procedures in Kim and Dieker (2011), Dieker and Kim (2012) avoid the Bonferonni inequality when comparing groups of three alternatives, but again requires the Bonferonni inequality for more than three alternatives. In this paper, we develop the Bayes-inspired IZ (BIZ) procedure, a fully sequential elimination procedure that satisfies the IZ guarantee (given assumptions on the sampling variances), is less conservative than existing IZ procedures, and samples less as a consequence. This procedure does Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 3 not use the Bonferroni inequality, instead using a novel symmetry based in Bayesian analysis. We assume independent normal samples, and present versions for both continuous and discrete time. The continuous-time BIZ procedure has a tight bound on worst-case preference-zone PCS: the PCS of the least-favorable configuration is exactly equal to the target probability. This is the first sequential elimination procedure with this property for more than two alternatives. The worst-case preference-zone PCS of the discrete-time BIZ procedure is shown in numerical experiments to be extremely close to the target PCS, even for as many as 214 = 16, 384 alternatives. The discrete-time BIZ procedure generalizes the non-elimination procedure PB⇤ introduced by Bechhofer et al. (1968), and the tight worst-case preference-zone PCS bound presented here also applies to a continuoustime version of PB⇤ . Our theoretical results (the IZ guarantee, and tightness of the worst-case preference-zone PCS bound) assume that the sampling variances are known, and are either common across alternatives, or are integer multiples of a common divisor. The BIZ procedure itself, however, allows both known and unknown sampling variances with arbitrary values, and numerical results suggest that the procedure’s performance is robust to deviations of the sampling variances from the structure assumed by the theoretical results. Although our bound on worst-case preference-zone PCS is tight in continuous time, and nearly tight in discrete time, BIZ’s PCS under configurations that are not least favorable can be strictly larger than the target. Thus, in these other configurations, BIZ also over-delivers on PCS. Furthermore, Wang and Kim (2012) shows that, for a variant of Paulson’s procedure, the contribution to over-delivery from the Bonferonni inequality is smaller than from the requirement that PCS be no smaller than the target for slippage configurations in the preference zone. However, violating this slippage configuration requirement would violate the IZ guarantee itself, while our results show that the Bonferonni inequality can be avoided while retaining the IZ guarantee. Numerical experiments demonstrate that, across a variety of configurations, BIZ’s over-delivery is much less than that of a leading IZ procedure, the KN procedure of Kim and Nelson (2001). The KN family of procedures “might be considered state-of-the-art for IZ R&S” (Branke et al. 2007), Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 4 and has been shown to be highly efficient compared to other existing IZ procedures (Malone et al. 2005, Wang and Kim 2012). Thanks to reduced over-delivery, BIZ requires fewer samples than KN on a variety of problems. Although the PCS bounds presented in this paper are non-Bayesian, the BIZ procedure is derived using a Bayesian approach. This derivation employs a Bayesian prior concentrated on slippage configurations. The proof techniques are reminiscent of results on the relationship between minimax and Bayesian analysis from decision theory (see, e.g., Berger (1985)). Thus, this work connects the IZ with the Bayesian formulation of R&S (see, e.g., Gupta and Miescke (1996), Chick and Inoue (2001), Frazier et al. (2008, 2009), Frazier and Powell (2008), Chick et al. (2010)). We begin in Section 2 by formally stating the IZ formulation of the R&S problem. We then introduce the BIZ procedure for discrete time in Section 3, first assuming a known common sampling variance in Section 3.1, and then the allowing sampling variances to be heterogeneous and unknown in Section 3.2. Our theoretical results, first that the BIZ procedure satisfies the IZ guarantee when the variance is known and is either common across alternatives or has an integer multiple structure, and second that it has tight worst-case preference-zone PCS bounds in continuous time, are given in Section 4. To support this analysis, a continuous-time generalization of the discrete-time BIZ procedure is also given. Section 5 gives numerical results, including a comparison with the KN procedure of Kim and Nelson (2001) and the PB⇤ procedure of Bechhofer et al. (1968). 2. Indi↵erence-Zone Ranking and Selection We have k alternative simulated systems, among which we would like to select the best. Samples from system x 2 {1, . . . , k } are normally distributed and independent, over time and across alternatives. Let µx and and = ( 1, . . . , pair µ, k) 2 x be the mean and variance of this sampling distribution. Let µ = (µ1 , . . . , µk ) be the corresponding vectors of sampling means and variances. Together, the are referred to as a system configuration. Our goal is to observe samples sequentially over time to find which alternative is the best, in the sense of having the largest µx . Let t = 0, 1, 2, . . . index time, and let Ytx be the sum of the samples observed from alternative x Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 5 2 x) by time t, so that Ytx is a discrete-time random walk with N (µx , increments and Y0x = 0. For any set A ✓ {1, . . . , k } let YtA be the vector (Ytx : x 2 A), and Yt be the vector Yt = (Yt1 , . . . , Ytk ). Any R&S procedure observes samples over time, either adaptively or deterministically, and either choosing to sample all of the alternatives at each time, or only a subset. Based on these samples, the procedure eventually stops sampling and selects an alternative as its estimate of the best. Call the selected alternative x̂. The goal in designing an R&S procedure is to take as few samples as possible, while still accurately selecting the best alternative. We now define the indi↵erence-zone guarantee, which is a statistical guarantee on the quality of the solution produced by an R&S procedure. First, we define the probability of correct selection as PCS(µ, ) = Pµ, ⇢ x̂ 2 arg max µx , x where Pµ, is the probability measure under which samples from system x have mean µx and variance 2 x. In the common-variance case, when 2 x = 2 for all x, we write PCS(µ, ) in place of PCS(µ, ) and Pµ, in place of Pµ, . Then, we define the preference zone (PZ), parameterized by PZ( ) = µ 2 Rk : µ[k] where µ[k] µ[k 1] ... µ[k 1] > 0, to be the set , µ[1] are the sorted components of µ. This is the set of system configura- tions under which the best alternative is better than the second best by at least . The complement of the preference zone is called the indi↵erence zone, and is the set of system configurations in which we are indi↵erent between the best and second best alternatives. Then, a procedure meets the indi↵erence-zone (IZ) guarantee at P ⇤ 2 (1/k, 1) and PCS(µ, ) P⇤ > 0 if for all µ 2 PZ( ). We assume that P ⇤ > 1/k because IZ guarantees for smaller values of P ⇤ can be met by choosing x̂ uniformly at random from among {1, . . . , k } without observing any samples. In this definition, whether a procedure meets the IZ guarantee depends upon , although procedures that satisfy the IZ guarantee are usually designed to do so for all 2 Rk++ . Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 6 3. The Bayes-inspired IZ (BIZ) Procedure In this section we define the Bayes-inspired IZ procedure (BIZ), and summarize the theoretical results shown later in Section 4. This procedure is developed using a Bayesian motivation, although the PCS bounds and IZ guarantee that we show are non-Bayesian. We first define a version for known common variance in Section 3.1, and then generalize to unknown and/or heterogeneous variances in Section 3.2. The BIZ procedure as described in this section operates in discrete time. In support of theoretical analysis, Section 4 provides generalizations of this discrete-time procedure that may operate in discrete or continuous time. 3.1. The BIZ Procedure with Known Common Variance We first define the Bayes-inspired indi↵erence zone (BIZ) procedure for the case of common known variance, when 2 x 2 are all equal to the constant . In Section 4, we show that this procedure satisfies the IZ guarantee, with tight bounds on worst-case preference-zone PCS in continuous time. Below, in Section 3.2, we generalize to heterogeneous and/or unknown sampling variances. BIZ is an elimination procedure. It maintains a list of alternatives that are in contention, and at each point in time, it takes one sample from each alternative in this set. Initially, all alternatives are in contention, and over time, as samples are observed, alternatives are eliminated. Once an alternative is eliminated, it may not come back into contention, and will not be considered for selection when sampling stops. When all but one alternative has been eliminated, this remaining alternative is selected as the best. In contrast with non-elimination procedures, which sample every alternative at every time, elimination procedures may eliminate bad alternatives quickly to reduce sampling e↵ort. BIZ is parameterized by the values P ⇤ 2 (1/k, 1) and > 0 for which we desire an IZ guarantee, and a parameter c satisfying c 2 [0, 1 (P ⇤ ) k c = 0 if k = 2. 1 1 ] if k > 2, (1) Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 7 The parameter c determines how aggressively we eliminate alternatives, and its choice is discussed below in Section 3.3. We recommend setting it to its maximum value, 1 2 x also depends on the sampling variances = 2 (P ⇤ ) k 1 1 . The procedure . For each t, x 2 {1, . . . , k }, and subset A ✓ {1, . . . , k }, we define a function qtx (A) = exp ✓ 2 Ytx ◆ X exp x0 2A ✓ 2 ◆ Ytx0 . (2) In Section 4, this expression is shown to be equal to a Bayesian posterior probability that alternative x is the best, given Yt , and given that the best is in the subset of alternatives A. The BIZ procedure for known common sampling variance is then defined by Alg. 1. Algorithm 1 BIZ for known common sampling variance, in discrete time Require: c 2 [0, 1 (P ⇤ ) k 1 {1, . . . , k }, t 1 ], > 0, P ⇤ 2 (1/k, 1), common sampling variance 1: Let A 0, P 2: while maxx2A qtx (A) < P do 3: while minx2A qtx (A) c do 4: Let x 2 arg minx2A qtx (A). 5: Let P 6: Remove x from A. P/(1 P ⇤ , Y0x 2 > 0. 0 for each x. qtx (A)). 7: end while 8: Sample from each x 2 A and add this sample to Ytx to obtain Yt+1,x . Then increment t. 9: 10: end while Select x̂ 2 arg maxx2A Ytx as our estimate of the best. The set A is the set of alternatives in contention, and is initially set to contain all of the alternatives in Step 1. Alternatives can be eliminated either in Step 6, or by the final selection in Step 10, which e↵ectively eliminates all remaining alternatives except the one selected. These eliminations are performed based on the current value of (qtx (A) : x 2 A), and an adaptively updated 8 Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 threshold P . The threshold P , which is initially set to P ⇤ , can be interpreted as a Bayesian posterior probability of selecting the best that we must achieve to stop sampling. Motivation: We motivate the BIZ procedure as follows. First, consider elimination resulting from exiting the outer “while” loop in Step 2 and going to Step 10. Recall that qtx (A) can be interpreted, in a Bayesian setting, as the posterior probability that x is the best (given that the best is in our contention set). The quantity P is a threshold on the probability of correctly selecting the best that we must achieve to stop sampling. In Step 2, if an alternative exceeds the this threshold P , then we exit the loop and select it as best in Step 10. Now, consider elimination resulting from entering the inner while loop in Step 3 and going to Step 7. The quantity minx2A qtx (A) is this posterior probability for the alternative in contention that is least likely to be the best. The inner while loop, Step 3, checks whether this minimal posterior probability is below the threshold c, and if it is, eliminates this alternative by removing it from A in steps 4 to 7. In addition to removing this alternative from A, the threshold P is increased, to account for the fact that we may have incorrectly removed the best alternative from A, and should strengthen the criteria that we must meet in Step 2 to stop. The behavior of this algorithm is illustrated in Figure 1. The example illustrated has k = 4 alternatives. The lower threshold c is plotted as a horizontal line, and the upper threshold P is plotted as a line with jumps. The posterior probability qtx (A) for each alternative x is also plotted versus time t. The figure uses the additional notation ⌧n to indicate the time at which the nth elimination occurs, and Zn to indicate the alternative eliminated. Starting from time t = 0, the contention set A contains all 4 alternatives and we plot qtx (A) for each. At time ⌧1 , minx2A qtx (A) hits the lower threshold c and the alternative Z1 achieving this minimum is eliminated. The values qtx (A) jump at this time, as an alternative is removed from A. This jump is small for the two alternatives with small posterior probabilities qtx (A), but is larger for the one alternative with a higher value. The threshold P also jumps. Moving forward from time ⌧1 , three alternatives remain in A, and we plot qtx (A) for each. A second alternative is eliminated at time ⌧2 when its posterior probability qtx (A) hits threshold c. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 9 Final selection 1 Pn 0.8 0.6 qtx(An) 0.4 0.2 c 0 0 Figure 1 time2000 (t) 8000 !6000 Z2) 2 (eliminate !1 (eliminate time (t) Z1) 4000 10000 Illustration of the BIZ procedure with k = 4 alternatives. The BIZ procedure follows the posterior probabilities qtx (A) over time, eliminating alternatives as they hit the lower threshold. Each time an alternative is eliminated, the upper threshold jumps upward. Eventually, an alternative reaches the upper threshold and is selected as the best. (At the same time, the largest posterior probability comes very close to the upper threshold, but does not hit it.) After time ⌧2 , posterior probabilities qtx (A) are plotted for the two remaining alternatives, until an alternative meets the upper threshold P (marked “Final selection”). At this time, the alternative whose qtx (A) hits the upper threshold is selected as best, and sampling stops. 3.2. The BIZ Procedure with Heterogeneous and Unknown Sampling Variances While Section 3.1 assumed a common known sampling variance 2 x = 2 , sampling variances are often heterogeneous and unknown in practice. In this section, we generalize the BIZ procedure to handle heterogeneous sampling variances, in both variance-known and variance-unknown settings. When the variances are known and are integer multiples of a common value, this procedure retains the IZ guarantee of the known common-variance BIZ procedure. The continuous-time version of this procedure presented in Section 4.5 also retains the IZ guarantee, with a tight worst-case preference-zone PCS bound. However, in discrete time, when the variances are unknown or lack an integer multiple structure, we do not have a proof that it satisfies the IZ guarantee. Instead, in this setting, we present this procedure as a heuristic. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 10 The discrete-time BIZ procedure for unknown and/or heterogeneous sampling variances is given below in Alg. 2. Rather than taking only one sample from each alternative in contention for each increment in t, Alg. 2 takes a variable number, storing the number of samples taken as ntx . We let Ztx = Yntx ,x be the sum of all of these samples. The algorithm also maintains an adaptive estimate b2tx of the sampling variance for alternative x, and is designed to keep ntx approximately proportional to b2tx . Alg. 2 accepts additional parameters beyond those accepted by Alg. 1: an integer n0 , and a collection of integers B1 , . . . , Bk . n0 is the number of samples to use in a first stage of samples, for which we recommend the value of 100. The parameter Bx governs the number of samples taken from alternative x in each stage. In practice we recommend setting Bx to 1, and we leave it as a free parameter because doing so supports theoretical analysis in Section 4.5. Alg. 2 can also be used when the sampling variances are known. In this case, we set n0 = 0 and replace the estimators b2tx in Alg. 2 with their known values. If the variances are known and identical, and we set B1 , . . . , Bk to their recommended values of 1, we recover Alg. 1. Alg. 2 uses the quantity qbtx (A), defined here in terms of another quantity qbtx (A) = exp ✓ Ztx t ntx ◆ X exp x0 2A ✓ Ztx0 t ntx0 ◆ , t P t. ntx0 0 = Px 2A . b2 x0 2A (3) tx0 Motivation We motivate Alg. 2 as follows. Consider what would happen if ntx were exactly proportional to b2tx and our estimate b2tx = 2 x were perfect, so ntx = 2 xt consider the stochastic process Y 0 = (Ytx0 : t = 0, 1, 2, . . .), where Ytx0 = Ztx / for some 2 x. > 0. Then A straightforward computation shows that this stochastic process is a random walk whose increments are normal with mean µx and variance 1/ . This variance does not depend on x, so to find arg maxx µx , we may use a common-variance R&S procedure, such as Alg. 1. Alg. 2 is derived by running Alg. 1 on Y 0 if these idealized conditions are met, or on an approximation to it if they are not. This approximation is Ytx0 ⇡ Ztx /( ntx / t ). Applying (2), but with this approximation of Ytx0 in place of Ytx and the variance 1/ of the increments of Ytx0 in place of 2 , provides (3). When, as in the motivating situation described above, b2tx = 2 tx and ntx = 2 tx t Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 11 Algorithm 2 Discrete-time implementation of BIZ, for unknown and/or heterogeneous variances. Require: c 2 [0, 1 (P ⇤ ) k 1 1 ], > 0, P ⇤ 2 (1/k, 1), n0 integers. Recommended choices are c = 1 sampling variances 1: 2 x (P ⇤ ) k 0 an integer, B1 , . . . , Bk strictly positive 1 1 , B1 = · · · = Bk = 1 and n0 = 100. If the are known, replace the estimators b2tx with the true values n0 = 0. To compute qbtx (A), use (3). For each x, sample alternative x n0 times and set n0x P ⇤, t {1, . . . , k }, P Let A 3: while maxx2A qbtx (A) < P do 4: 5: 6: 7: 8: 9: 10: 11: and set n0 . Let Z0x and b20x be the sample mean and sample variance respectively of these samples. Let t 2: 2 x, 0. 1. while minx2A qbtx (A) c do Let x 2 arg minx2A qbtx (A). Let P P/(1 qbtx (A)). Remove x from A. end while Let z 2 arg minx2A ntx /b2tx . ⇣ ⌘ For each x 2 A, let nt+1,x = ceil b2tx (ntz + Bz )/b2tz . For each x 2 A, if nt+1,x > ntx , take nt+1,x ntx additional samples from alternative x. Let Zt+1,x and b2t+1,x be the sample mean and sample variance respectively of all samples from alternative x thus far. 12: Increment t. 13: end while 14: Select x̂ 2 arg maxx2A Ztx /ntx as our estimate of the best. for some > 0, we have t = t = ntx / 2 x, terms cancel, and Ytx0 is exactly equal to its approx- imation. Below, in Section 4.5, we analyze special cases in which this occurs and show that, in these situations, Alg. 2 satisfies the IZ guarantee, and does so with a tight bound on worst-case preference-zone PCS in continuous time. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 12 3.3. BIZ Recovers PB⇤ as a Special Case When c = 0 and variances are known and common, the discrete-time BIZ procedure is equivalent to the non-elimination procedure PB⇤ introduced by Bechhofer et al. (1968). This can be seen as follows: Because Ytx is almost surely finite at any fixed t, qtx (A0 ) > 0 almost surely. Thus, in Step 3 of Alg. 1, minx2A qtx (A) > 0 = c, the loop from Steps 4 to 7 will never execute, and A = {1, . . . , k } and P = P ⇤ . The resulting procedure is then no longer an elimination procedure, and takes one sample from every alternative at each point in time. It stops and selects the alternative with the largest sample mean at the first time t for which maxx=1,...,k qtx ({1, . . . , k }) P ⇤ . This is exactly the PB⇤ procedure of Bechhofer et al. (1968). The parameter c determines the trade-o↵ between the number of stages of sampling and the overall number of samples taken. When c = 0 the procedure does no elimination. When c > 0, the procedure eliminates alternatives, with larger c causing more aggressive elimination, decreasing the number of samples and increasing the number of stages. In highly parallel computing environments, and some biological and agricultural applications, one can evaluate many alternatives simultaneously, and the focus is on minimizing the number of stages. In simulation, however, when the number of alternatives is large compared to the parallelism in one’s computing environment, the focus is on minimizing the number of samples taken. In Section 5 we compare the number of samples taken by BIZ with c at its maximum value, 1 (P ⇤ ) k c=1 1 1 , to the number taken by PB⇤ (which is BIZ with c at its minimum value, 0). Setting (P ⇤ ) k 1 1 dramatically reduces the expected number of samples taken, particularly when some alternatives are much worse than others. For use in simulation in a non-parallel setting, we recommend c = 1 (P ⇤ ) k 1 1 . 4. Theoretical Analysis In this section, we present our theoretical results: that IZ guarantees hold for Alg. 1 and, if variances are known and have a special structure, Alg. 2; and that continuous-time generalizations of these procedures have tight bounds on worst-case preference-zone PCS. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 13 First, we present a generalization of the R&S problem that allows both continuous-time and discrete-time sampling in Section 4.1. Then, we consider the setting with common known variance: Section 4.2 generalizes Alg. 1 to the continuous-time setting; Section 4.3 presents preliminary definitions and results; and Section 4.4 presents the main theoretical results for the common known variance setting. Section 4.5 considers the setting with heterogeneous known variance, presenting first a continuous-time analogue of Alg. 2, and then theoretical results for both this continuous-time analogue and Alg. 2 itself. 4.1. Generalization to Continuous Time: Observation Process Although the R&S problem occurs in discrete time in practice, our theoretical analysis relies on a generalization in which observations occur in continuous time, while decisions about eliminating alternatives or stopping to select the best are made at any of a predetermined set of decision points. When this set of decision points is the non-negative integers Z+ = {0, 1, . . .}, we recover the discretetime BIZ procedure. When it is the non-negative reals R+ = [0, 1), we obtain a continuous-time version of BIZ, which we later show has tight worst-case bounds on preference-zone PCS. Recall from Section 2 that, in discrete time and under Pµ, , the sum of all samples from alternative x by time t, given in the stochastic process (Ytx : t 2 Z+ ), is a discrete-time random walk with N (µx , 2 x) increments. We generalize this by letting (Ytx : t 2 R+ ) be a Brownian motion under Pµ, starting from 0, with drift µx , volatility x, and independence across x. This is consistent with the previous definition of Ytx at integer times t, since (Ytx : t 2 Z+ ) continues to be a discrete-time random walk with N (µx , 2 x) increments. As before, for each A ✓ {1, . . . , k } we let YtA = (Ytx : x 2 A), and Yt = (Yt1 , . . . , Ytk ). We let F = (Ft : t 2 R+ ) be the filtration generated by (Yt : t 2 R+ ). In the continuous-time setting, we assume that the variances are known. If they are not known, they can be estimated with perfect accuracy from a sample path (Ytx : 0 t ✏) for any ✏ > 0. 4.2. Generalization to Continuous Time: The BIZ Procedure with Common Variance We now generalize the BIZ procedure for known common variance in discrete time (Alg. 1) to include the continuous-time setting. In this section, we assume 2 x = 2 for all x, with 2 known. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 14 This generalized BIZ procedure includes a parameter T which is a set of decision points, and is set to either R+ or Z+ . Setting T = R+ provides a continuous-time procedure, while setting T = Z+ recovers the discrete-time procedure Alg. 1. To define this generalized BIZ procedure, we recursively define a sequence of stopping times 0 = ⌧ 0 ⌧ 1 · · · ⌧k 1 1, random variables Z1 , . . . , Zk 1 and P0 , P1 , . . . , Pk 1 , and random sets A0 , A1 , . . . , Ak 1 . We first define ⌧0 , A0 and P0 as P0 = P ⇤ , ⌧0 = 0, Then, for each n = 0, 1, . . . , k A0 = {1, . . . , k }. (4a) 2, we define ⌧n+1 , Zn+1 , An+1 and Pn+1 recursively given ⌧n , An , and Pn as ⇢ ⌧n+1 = inf t 2 T \ [⌧n , 1) : min qtx (An ) c or max qtx (An ) x2An x2An Zn+1 2 arg min q⌧n+1 ,x (An ), x2An Pn , (4b) An+1 = An \ {Zn+1 }, ✓ ◆ Pn+1 = Pn 1 min q⌧n+1 ,x (An ) . x2An Finally, with these quantities defined, x̂ is the single alternative in Ak 1 , x̂ 2 Ak 1 . In this definition, the times ⌧1 , . . . , ⌧k variables Z1 , . . . , Zk 1 1 (4c) are times at which alternatives are eliminated, the random are the alternatives eliminated at these times, and An is the set of alternatives in contention starting at time ⌧n . The random variables P0 , . . . , Pk 1 are thresholds our posterior probability of being best must achieve to allow us to stop sampling. The times at which alternatives are eliminated has a particular structure: Initially, elimination occurs because minx2An qtx (An ) c, i.e., because this posterior probability of being best fell below the lower threshold c. Eventually though, an elimination occurs because maxx2An qtx (An ) Pn , i.e., because an alternative’s posterior probability of being best exceeded the upper threshold Pn . At this time, Lemma 6 below shows that all alternatives except one are eliminated simultaneously, ⌧n+1 = ⌧n+2 = · · · = ⌧k 1 , and that Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 the one alternative remaining in Ak qtx (An ) 1 15 (which is selected as best) is the one whose qtx (An ) satisfied Pn at time t = ⌧n+1 . We define a random variable M so that the time at which this occurs is ⌧M = ⌧n+1 , and the M th through the (k 1)st eliminations occur simultaneously in this way. When T = Z+ , the algorithm defined by (4) is identical to Alg. 1. We see this as follows. The stopping time ⌧n is the value of t in Alg. 1 at the time when the nth alternative is eliminated, either explicitly in Step 6 (if n < M ), or implicitly in Step 10 (if n M ). For 1 n < M , at time ⌧n , we go through the inner while loop (Steps 2 through 7) for the nth time. The alternative x chosen for elimination in Step 4 is Zn , Step 5 takes P from Pn 1 to Pn , and Step 6 takes A from An 1 to An . At time t = ⌧M , the condition of the outer while loop checked in Step 2 fails to be satisfied, and Alg. 1 goes to Step 10 and selects arg maxx2A Ytx = arg maxx2A Y⌧M ,x , which Lemma 6 below shows is the same as the selection x̂ made by the procedure defined by (4). In (4b), the choice among the arg min set for Zn+1 does not a↵ect the analysis, but we set Zn+1 to be the alternative with the smallest index in that set. When ⌧n+1 = 1, we choose Pn+1 = 1 and Zn+1 uniformly at random from An , although again this choice does not a↵ect the analysis because later in Lemma 4 we show that ⌧n+1 < 1 almost surely under any Pµ, . The claim that each ⌧n is a stopping time is justified in Section 4.3, in Lemma 5. This definition of BIZ in continuous time also provides an extension of PB⇤ to continuous time, obtained by setting c = 0. The resulting procedure can be simplified, as is shown below in Lemma 7, to a procedure that samples from all alternatives, eliminating none, until a selection is made at time ⌧ = ⌧1 = ⌧2 = · · · = ⌧k ⌧ = inf 1 ( of x̂ 2 arg maxx=1,...,k Y⌧,x . This time can be written ) k X 2 ⇤ 2 t 2 T : max exp Ytx / P exp Ytx0 / . x x0 =1 When T = Z+ , this is the original discrete-time PB⇤ procedure of Bechhofer et al. (1968). 4.3. Preliminaries for the Proofs In this section, we present definitions and preliminary results that support our main theoretical results for the common variance setting in Section 4.4. We continue to assume that x, with 2 known. 2 x = 2 for all Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 16 We first construct a probability measure Q under which the vector of sampling means is chosen at random according to a prior distribution. To indicate that the vector of sampling means under Q is random, and may di↵er from the true vector of sampling means µ, we denote it by ✓. We emphasize that Q is a mathematical construct that we use to analyze the BIZ procedure, and is di↵erent from the true sampling distribution Pµ, . We construct Q as follows. Let X ⇤ be chosen uniformly at random from among 1, . . . , k, and let ✓X ⇤ = . Let ✓x = 0 for all x 6= X ⇤ . Configurations of the form ✓X ⇤ ✓x = for some parameter literature, and slippage configurations in which > 0 are called slippage configurations in the R&S = are often the most difficult configurations under which to select correctly. We then define a family of probability measures that includes and generalizes Q. For each u 2 Rk with u[k] 6= u[k 1] , define a probability measure Qu as follows. First, let (R(1), . . . , R(k)) under Qu be a uniformly distributed permutation of (1, . . . , k). Then, let ✓x = uR(x) almost surely under Qu , and let X ⇤ 2 arg maxx ✓x . (This argmax is unique because u[k] 6= u[k 1] .) Given ✓, we let each (Ytx : t 2 R+ ) be an independent Brownian motion under Qu with drift ✓x and volatility . Defining u = [ , 0, . . . , 0], we have Q = Qu , so this definition generalizes the previously defined Q. The following lemma provides an expression for the posterior probability that a specified alternative x0 has the largest sampling mean, given the prior Qu and partial information about the permutation R. Proofs of this and other lemmas may be found in the appendix. Lemma 1. Suppose 2 x = 2 > 0 8x. Let A ✓ {1, . . . , k } with x0 2 A. Let u 2 PZ( ) and r⇤ 2 arg maxx ux . Let R be the (random) set of permutations r with R(x) = r(x) for x 2 / A. Then, ⇤ 0 Qu {X = x | YtA , (R(x))x2A / }= X r2R:r(x0 )=r⇤ exp 1 X 2 x2A Ytx ur(x) ! X r2R exp 1 X 2 ! Ytx ur(x) . x2A If R contains no permutations r with r(x0 ) = r⇤ , then the numerator in this expression is 0. Lemma 1 has as a consequence Lemma 2 below, which gives the posterior probability under Q that alternative x has the best sampling mean, given that the best is in a specified set A. This expression is exactly qtx (A) (with x0 in place of x), defined earlier in (2). As discussed in Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 17 Section 3.1, this interpretation of qtx (A) motivates the BIZ procedure, and as we will see later, plays an important role in its analysis. Lemma 2. Suppose 2 x = 2 ⇤ > 0 8x. Let A ✓ {1, . . . , k } and x0 2 A. Then, 0 ⇤ Q {X = x | YtA , X 2 A} = exp ✓ 2 Ytx0 ◆ X exp x2A ✓ 2 ◆ Ytx . The expression is una↵ected if we also condition on (R(x))x2A / . Later, we also use the following monotonicity result. Its proof involves algebraic manipulations of expressions from Lemmas 1 and 2. Lemma 3. Suppose 2 x = 2 > 0 8x. Fix u 2 PZ( ), a permutation r0 of the integers {1, . . . , k }, and y 2 Rk . Let A be a non-empty subset of {1, . . . , k }. Let B denote the event that X ⇤ 2 A, Ytx = yx for each x 2 A, and R(x) = r0 (x) for each x 2 / A. Then max Qu {X ⇤ = x | B } max Q {X ⇤ = x | B } , (5) min Qu {X ⇤ = x | B } min Q {X ⇤ = x | B } . (6) x2A x2A x2A x2A Our results also require the following pair of technical lemmas. The first states that, with probability 1, the BIZ procedure takes finitely many samples. Its proof employs a standard geometric decay argument. The second states that the elimination times are stopping times of the filtration generated by the observation process Y , and uses elementary manipulations of events. Lemma 4. Suppose 2 x = 2 > 0 8x. Then ⌧n < 1 a.s. under Pµ, for n = 0, 1, . . . , k Lemma 5. For each n = 0, 1, . . . , k 1 and µ 2 Rk . 1, ⌧n defined by (4) is a stopping time of F . In Section 4.2, we stated that, at the first elimination time ⌧n+1 caused by an alternative’s qtx (An ) exceeding Pn , all other alternatives are eliminated simultaneously, and this alternative is selected as the best. We describe this behavior more formally in the following lemma, whose statement uses the definition of the random variable M , ⇢ M = inf n = 1, . . . , k 1 : max q⌧n ,x (An 1 ) x2An 1 Pn 1 , Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 18 so that ⌧M is the first time at which we eliminate an alternative because maxx2An 1 q⌧n ,x (An 1 ) exceeds Pn 1 . 2 x Lemma 6. Suppose 2 = > 0 8x. Then, for any µ 2 Rk , the following statements hold almost surely under Pµ, . (a) M k 1. M and x̂ 2 arg maxx2AM (b) ⌧n = ⌧M for all n (c) If T = R+ , ⌧M 1 1 Y⌧M ,x . < ⌧M . Lemma 6 allows us to formally state the previously claimed simplification of BIZ when c = 0, from the discussion in Section 4.2 of the PB⇤ procedure. Lemma 7. When c = 0, we have ⌧1 = ⌧2 = · · · = ⌧k ⌧ = inf ( t 2 T : max exp x Ytx / 1 = ⌧ and x̂ 2 arg maxx=1,...,k Y⌧,x , where 2 P ⇤ k X exp Ytx0 / x0 =1 2 ) . We may now state two lemmas, which together constitute the proof of the main result in Section 4.4. These lemmas use CS = {x̂ 2 arg maxx ✓x , ⌧k 1 < 1} to denote the event of correct selec- tion. Lemma 8 shows that the non-Bayesian probability of correct selection PCS(u, ) is identical to the probability of correct selection under the Bayesian prior Qu . The proof follows a symmetry or “equalizing” argument. Lemma 8. Suppose 2 x 2 = > 0 8x. Let u 2 Rk . Then Qu {CS} = PCS(u, ). Furthermore, PCS(u, ) is invariant to translations and permutations of u. Lemma 9 shows that the conditional Bayesian PCS is bounded below by the random variable Pn 1 , with equality in the case of continuous-time sampling and prior Q. Lemma 9. Suppose 2 x = 2 > 0 8x. Then, for each n = 1, . . . , k ⇤ Qu CS | F⌧n , (R(x))x2A / n 1 , X 2 An 1 , M 1 and each u 2 PZ( ), n If T = R+ and u = [ , 0, . . . , 0] then this inequality holds with equality. Pn 1 . Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 19 4.4. Theoretical Results for the Common Variance Setting Using the preliminary results from the previous section, we now state and prove our main result common known variances, Theorem 1. The first statement shows that the BIZ procedure satisfies the IZ guarantee in both discrete and continuous time. The second statement in the theorem shows that, in continuous time, the bound on the worst-case preference-zone PCS of the BIZ procedure is tight. This second statement can be interpreted as showing that any slack in the PCS bound for BIZ in discrete time is due to the gap between the time at which the continuous-time procedure would eliminate an alternative and the next integer-valued time. Theorem 1. Let c 2 [0, 1 (P ⇤ ) k 1 1 ], T 2 {Z+ , R+ }, > 0, P ⇤ 2 (1/k, 1), and 2 x = 2 > 0 for all x. Then, under the BIZ procedure defined by (4), P⇤ PCS(µ, ) for all µ 2 PZ( ). Furthermore, if T = R+ , inf µ2PZ( ) Proof: PCS(µ, ) = P ⇤ . Let µ 2 PZ( ). Let µ0 be a permutation of µ such that µ01 given by ux = µ0x µ01 + . Thus, u1 = µ0x for all x. Let u 2 Rk be and ux 0 for all x 6= 1. Because u is a permutation and translation of µ, Lemma 8 implies PCS(µ, ) = Qu {CS}. Furthermore, u 2 PZ( ). Since X ⇤ 2 A0 and M (7) 1 with probability 1, and the complement of A0 is empty, taking n = 1 in Lemma 9 shows Qu {CS | F⌧1 } = Qu {CS | F⌧1 , X ⇤ 2 A0 , M 1} P0 = P ⇤ . (8) Then, the tower property of conditional expectation provides Qu {CS} = EQu [Qu {CS | F⌧1 }] EQu [P ⇤ ] = P ⇤ , where EQu is the expectation under Qu . Combining (7) and (9) provides PCS(µ, ) (9) P ⇤. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 20 We have shown that PCS(µ, ) P ⇤ for all µ 2 PZ( ). This shows that inf µ2PZ( ) PCS(µ, ) P ⇤. To see that the infimum is equal to P ⇤ when T = R+ , consider µ = u = [ , 0, . . . , 0]. Lemma 9 shows that, in this case, the inequalities in (8) and (9) are actually equalities. Combining equality in (9) with the equality (7) shows that PCS([ , 0, . . . , 0], ) = P ⇤ , implying inf µ2PZ( ) PCS(µ, ) P ⇤ . This shows that the infimum must in fact equal P ⇤ . ⇤ The last paragraph of the proof shows that the infimum of the PCS over the preference zone is achieved by the configuration [ , 0, . . . , 0]. The invariance of the PCS to translations and permutations of the configuration shown by Lemma 8 implies that the infimum is also attained by any slippage configuration with parameter . Thus, these slippage configurations are least-favorable for the BIZ procedure with common known variance. 4.5. Theoretical Results for the Heterogeneous Variance Setting We now discuss the heterogeneous known variance setting. We present a continuous-time procedure that is analogous to the discrete-time Alg. 2. This continuous-time procedure satisfies the IZ guarantee and has tight worst-case preference-zone PCS bounds. We then use this fact to show that Alg. 2 satisfies the IZ guarantee when variances are known and have a special integer multiple structure. For each x let nx (t) = 2 x t. This quantity is the continuous-time analogue of the discrete- time quantity ntx in Alg. 2, and in the certain special cases discussed below, nx (t) = ntx for all integer times t. Now define a stochastic process (Ytx0 : t 0) as Ytx0 = Ynx (t),x / 2 x. A straightforward computation shows that (Ytx0 : t 2 R+ ) is a Brownian motion with drift µx and volatility 1/ , so any algorithm that performs R&S in the continuous-time common-variance case can be run on the modified observation processes Y 0 , and the result is a continuous-time R&S algorithm for the original observation process Y . This was also noted for discrete time in Section 3.2. With this motivation, the continuous-time BIZ procedure for known heterogeneous variances is obtained by applying the continuous-time BIZ procedure for common sampling variances from (4) to the modified observation process Y 0 . More explicitly, this procedure is defined by first setting Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 P0 = P ⇤ , ⌧0 = 0, 21 A0 = {1, . . . , k }, (10a) then defining recursively, for n = 0, 1, . . . , k 2, ⇢ 0 0 ⌧n+1 = inf t 2 T \ [⌧n , 1) : min qtx (An ) c or max qtx (An ) x2An x2An Pn , Zn+1 2 arg min q⌧0 n+1 ,x (An ), x2An (10b) An+1 = An \ {Zn+1 }, ✓ ◆ 0 Pn+1 = Pn 1 min q⌧n+1 ,x (An ) . x2An 0 where qt,x (A) is obtained by substituting Y 0 in place of Y and 1/ in place of 0 qt,x (A) = exp ( Ytx0 ) X exp ( Ytx0 0 ) = exp x0 2A ✓ 2 x Ynx (t),x ◆ X exp x0 2A ✓ 2 x0 2 Y , n0x (t),x0 ◆ , (10c) and finally letting the selected alternative x̂ be the single entry in Ak 1 , x̂ 2 Ak 1 . When sampling variances are identically equal to (10d) 2 across alternatives and = 1/ 2 , we have 0 nx (t) = t, Ytx0 = Yt,x , qt,x (A) = qt,x (A), and the procedure defined by (10) is identical to (4). The following theorem shows that this procedure satisfies the IZ guarantee, and its worst-case preference-zone PCS bound is tight in continuous time. This is true even when sampling variances di↵er from each other. Theorem 2. Let c 2 [0, 1 (P ⇤ ) k 1 1 ], T = {Z+ , R+ }, > 0, P ⇤ 2 (1/k, 1), and 2 1 > 0, . . . , 2 k > 0. Then, under the BIZ procedure defined by (10), PCS(µ, ) P⇤ for all µ 2 PZ( ). Furthermore, if T = R+ , inf PCS(µ, ) = P ⇤ . µ2PZ( ) Proof: and Let x̂ be the selection decision x̂ defined by (10), with the specified values for P ⇤ , ,c,T, 2 1, . . . , 2 k. We use the superscript to emphasize that x̂ assumes sampling variances 2 x. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 22 Let x̂ be the selection decision defined by (4) using 2 = 1/ , and the same specified values of P ⇤ , , c, and T. The distribution of Y 0 under Pµ, is equal to the distribution of Y under Pµ, . Consequently, the distribution of x̂ under Pµ, is equal to the distribution of x̂ under Pµ, . This implies that Pµ, x̂ 2 arg maxx µx = Pµ, {x̂ 2 arg maxx µx }. The result then follows from applying Theorem 1 to Pµ, {x̂ 2 arg maxx µx }. ⇤ While (10) is directly implementable in continuous time, it is more difficult to apply in discrete time. While one can set T to Z+ in (10), the resulting procedure is not always implementable in discrete time. The reason is that (10) requires observations of Ynx (t),x for t 2 T. If nx (t) = 2 xt can fail to be an integer for some t 2 T, then these observations may be unavailable in discrete time. However, if the variances have a special integer multiple structure, then (10) is implementable in discrete time, and is equivalent to Alg. 2. In particular, suppose the variances satisfy 2 x = ax 2 for some common 2 and integers a1 , a2 , . . . , ak . If we set 2 x = 1/ are known and 2 and T = Z+ , then nx (t) = ax t is always an integer for t 2 Z+ , and all observations of Ynx (t),x required by (10) are available in discrete time. Furthermore, in this case, (10) is identical to Alg. 2 with parameters Bx = ax , n0 = 0, and b2x = 2 x. A direct consequence of this and Theorem 2 is that Alg. 2 satisfies the IZ guarantee, in this special case. We have just shown the following corollary to Theorem 2. Corollary 1. Let 2 x = ax 2 for all x, where ax 2 Z+ with ax 1, 2 > 0. Let c 2 [0, 1 (P ⇤ ) k 1 1 ], > 0, and P ⇤ 2 (1/k, 1). Then, under the BIZ procedure for known heterogeneous sampling variances given in Alg. 2 with n0 = 0, Bx = ax , and b2tx = PCS(µ, ) P⇤ 2 x for all x, for all µ 2 PZ( ). Outside of the common variance setting, the integer multiple structure assumed by Corollary 1 is unlikely to appear in practice. Also, in practice one would set Bx to 1, rather than to the values assumed by Corollary 1, to improve the responsiveness of the algorithm and reduce expected sample sizes. Thus, while Corollary 1 provides insight into the behavior of Alg. 2, it is not designed to provide an IZ guarantee that directly applies to how this algorithm is used in practice. Instead, we Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 23 present Alg. 2 as a heuristic in practical settings, and we use numerical experiments to investigate its statistical properties in the next section. 5. Numerical Results We demonstrate the performance of the BIZ procedure in discrete time with maximum elimination (c = 1 (P ⇤ ) k 1 1 ) on standard test problems, and compare it to another leading IZ procedure, the KN procedure of Kim and Nelson (2001), first on problems with common known sampling variance, then on problems with common unknown sampling variance, and finally on problems with heterogeneous unknown sampling variance. The KN procedure improves over previously proposed IZ procedures in a number of configurations (Kim and Nelson 2001), and the KN family of procedures has been regarded by Kim and Nelson (2006) and Branke et al. (2007) as state-of-the-art for IZ R&S. Improvements of the original KN procedure, particularly the KN++ procedure of Goldsman et al. (2002) and the KVP and UVP procedures of Hong (2006), o↵er better performance than KN in some settings with unknown and/or heterogeneous sampling variance, but KN remains one of the best existing IZ procedures. In problems with known variance, we modify KN from its original version in Kim and Nelson (2001) to take advantage of knowing the variance. Where the original procedure uses estimates of the variance, the modified procedure uses the actual value. The modified procedure also uses the tighter constant (h⇤ )2 = 2c⌘ ⇤ in place of the parameter h2 from Kim and Nelson (2001), where c = 1, and ⌘ ⇤ satisfies g(⌘ ⇤ ) = 12 exp( ⌘ ⇤ ) = 1 P⇤ . k 1 We set n0 = 1. This modified procedure is the same as the P procedure in Wang and Kim (2012), and when used in common variance configurations, the same as Paulson’s procedure Paulson (1964). In problems where the sampling variance is unknown, we use KN as originally described in Kim and Nelson (2001), with c = 1 and n0 = 100. We also compare to the PB⇤ procedure of Bechhofer et al. (1968), which is BIZ with no elimination, as described in Section 3.3. In our figures, we denote the PB⇤ procedure by BKS, the initials of the authors of Bechhofer et al. (1968). We examine both the PCS and the expected total number of samples taken, denoted E[N ]. We emphasize that N counts the total number of samples taken, and so a procedure without elimination Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 24 has E[N ] = kE[⌧k 1 ], while a procedure with elimination has E[N ] kE[⌧k 1 ]. Rather than plotting E[N ] directly, we plot the expected number of samples taken divided by the number of alternatives, E[N ]/k. This normalizes E[N ] and clarifies performance trends. Figure 2 shows the performance of KN, BKS, and BIZ (Algorithm 1), under three di↵erent configurations with common known variance, described in more detail below. Each row shows performance under a di↵erent configuration vs. the number of alternatives k. Left-hand panels show E[N ]/k and right-hand panels show PCS, obtained using 10,000 independent replications. SC: Row 1 of Figure 2 shows performance under a slippage configuration (SC), in which µ1 = and µx = 0 for x 6= 1. For many procedures, including BIZ and KN, a slippage configuration with parameter is the configuration in the preference zone in which correctly selecting the best is most difficult, and is often used as a test case to better understand the behavior of R&S procedures. Here, = 1, x = 2 = 100, and P ⇤ = 0.9. Experiments were performed at k = 2, 3, . . . , 8, and then at integral powers of 2 up to k = 214 = 16, 384. PCS under this SC is shown in the right-hand panel of Row 1. We know from their IZ guarantees that PCS for all three procedures is bounded below by the target probability P ⇤ = 0.9. Apparent deviations below 0.9 are due to estimation error — standard errors for PCS reported for BIZ and p BKS are approximately 0.9 ⇥ 0.1/104 = .003. Under KN, which has a loose worst-case preferencezone PCS bound and which over-delivers on PCS for large problems, PCS quickly rises away from P ⇤ as the number of alternatives grows. In contrast, under BIZ and BKS, PCS remains close to P ⇤ . The proximity of PCS to the target shows that, although the lower bound on worst-case preference-zone PCS given in Theorem 1 is no longer tight as we move from continuous to discrete time, the bound remains nearly tight in discrete time, at least in the settings tested. E[N ]/k under this SC is shown in the left-hand panel of Row 1. Points plotted have standard error less than 2 for KN and BIZ, and less than 7 for BKS. As the number of alternatives grows large, KN begins taking a very large number of samples, while the number of samples taken by BIZ grows at a much slower rate. For the largest problem considered, k = 16, 384, KN takes (22.4 ± 0.2) ⇥ 106 Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 1 KN BKS 1500 BIZ 2000 0.98 0.96 PCS E[N]/k 25 1000 KN BKS BIZ 0.94 0.92 500 0.9 0 0.88 1 10 100 1000 10000 1 300 1 250 0.98 200 100 0.94 0.9 0 0.88 1 10 100 1000 10000 1 Number of alternatives (k) 2200 2000 1800 1600 1400 1200 1000 800 600 400 200 0 10 100 1000 10000 Number of alternatives (k) 1 KN BIZ 0.95 PCS E[N]/k 10000 0.92 50 0.9 KN BIZ 0.85 0.8 0.75 1 10 100 1000 10000 1 Number of alternatives (k) Figure 2 1000 KN BKS BIZ 0.96 KN BKS BIZ 150 100 Number of alternatives (k) PCS E[N]/k Number of alternatives (k) 10 10 100 1000 10000 Number of alternatives (k) Common known variance: Rows shows PCS and E[N ]/k vs. the number of alternatives k under the following configurations: SC (row 1); MDM (row 2); and RPI (row 3). For SC and MDM, P ⇤ = 0.9. For RPI, P ⇤ = 0.8. Sampling variances are common and known. samples in expectation, while BIZ takes (7.574 ± .003) ⇥ 106 . The number of samples taken by KN is 3 times larger than the number taken by BIZ. Although both BIZ and BKS deliver a PCS that is close to the target, BIZ requires many fewer samples than BKS (a factor of 4.4 times fewer samples at k = 16, 384) because of its ability to eliminate poor alternatives early. This ability and its associated improvement in sampling efficiency, while clearly present, is relatively modest here in this slippage configuration where all of the Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 26 suboptimal alternatives have the same true sampling mean. In the configuration to be examined next, the ability to eliminate alternatives plays a much larger role. MDM: Row 2 of Figure 2 shows performance under the monotone decreasing means (MDM) configuration, in which µx = x. Unlike a slippage configuration in which the best is exactly better than the second best, where it is difficult to tell the best alternative from the others, the MDM configuration allows easier identification of the best alternative. As with the SC considered previously, = 1, x = 2 = 100, P ⇤ = 0.9, and experiments were performed at k = 1, 2, 3, . . . , 8, and then at integral powers of 2 up to k = 214 = 16, 384. In the right-hand panel of Row 2, we again see that KN’s PCS grows quickly away from the target, and is indistinguishable from 1 for k > 100 on the scale at which the figure is plotted. In contrast, BKS’s PCS stays close to the target. BIZ’s PCS is further from the target for intermediate values of k than it was under the SC, which shows that BIZ over-delivers on PCS when the configuration is not least favorable. However, BIZ’s over-delivery is not extreme, as its PCS stays below 0.99 over all k, and comes back to the target for large values of k. The left-hand panel of Row 2 shows E[N ]/k in the MDM configuration. Points plotted have standard error less than 2 in all cases (shrinking to less than 0.02 for k 512 for BIZ and KN). Here, as in the SC, BIZ outperforms KN across all values of k, although by a smaller margin. At k=16,384, KN requires 34, 242 ± 17 samples in expectation, 1.5 times more than the estimated 22,170 required by BIZ (the standard error on BIZ’s estimated sample size is less than .001 ⇥ 16, 384 = 16.4). BIZ also performs significantly better than BKS across all values of k. These aspects of the behavior under the MDM configuration are similar to those seen under the SC. The left-hand panel of Row 2 also shows a number of di↵erences between MDM and SC. Because the alternatives that are added to the MDM configuration as k increases are progressively further from the best, an MDM configuration with 10, 000 alternatives is not much more difficult than one with 100 alternatives. Thus, as the number of alternatives grows large in the MDM configuration, the number of samples E[N ] taken by a good procedure should grow very slowly. This is the reason why the total number of samples required per alternative, E[N ]/k, which initially rises under all Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 27 three procedures, eventually decreases under BIZ and KN. In contrast, the average number of samples required per alternative under BKS reaches a plateau and does not drop, causing the sampling e↵ort to grow linearly in k. This is because BKS is unable to eliminate bad alternatives, and so must sample all of them until the stopping time ⌧ . Beyond approximately 10 alternatives, adding more alternatives does not increase E[⌧ ] under BKS, but it does significantly increase E[N ] because E[⌧ ] = kE[N ]. This demonstrates the importance of elimination. RPI: Row 3 of Figure 2 shows performance of KN and BIZ on random problem instances (RPI). To generate a single random problem instance, we first choose k = ceil(exp(U )), where U is uniform between 0 and log(16, 384), and ceil(x) is the smallest integer greater than or equal to x. When generated in this way, k 2 with probability 1. We then generated the sampling mean of each of the k alternatives randomly from an independent normal distribution with mean 0 and variance 0.25. Those sampling means that are not best, but are within to be exactly 2 = 100, of the best, are then rounded down from the best. This ensures that the configuration is in the preference zone. Here, = 0.5, and P ⇤ = 0.8. We generated 75 random problem configurations in this way. The right-hand panel of Row 3 shows that PCS under both procedures for each problem instance is at or above P ⇤ = 0.8, as predicted by the theory. (There is one configuration with an estimated BIZ PCS of 0.797 ± .004, which is below 0.8, but not by a statistically significant margin.) In this RPI setting (as in the SC and MDM settings), KN’s PCS is larger than that of the BIZ procedure, and it becomes extremely close to 1 as k increases. On several large problem configurations, KN selected correctly on every one of the 10, 000 independent replications. This indicates extreme over-delivery. In contrast, BIZ’s PCS is lower than that of KN, and is evenly distributed within the interval [P ⇤ , 1] = [0.8, 1] for small k. As k increases, BIZ’s typical PCS moves toward 1, with typical estimated PCS for the largest problems near 0.995. All estimated BIZ PCS values are less than 0.997. While closer to 1 than under SC and MDM, BIZ’s over-delivery on PCS under these random problem instances is still less than that of KN. The left-hand panel of Row 3 shows E[N ]/k vs. k for both BIZ and KN. Standard errors are all less than 5. BIZ requires many fewer samples than KN, especially for large problems. As k grows, Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 28 KN-UNK BIZ-UNK 2000 PCS E[N]/k 1500 1000 500 0 1 10 100 1000 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 10000 KN-UNK BIZ-UNK 1 Number of alternatives (k) 300 PCS E[N]/k 200 150 100 50 0 1 10 100 1000 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 10000 1 100 1000 10000 0.95 1500 PCS E[N]/k 10 1 KN-UNK 2000 BIZ-UNK 1000 500 0.9 0.85 0.8 0 KN-UNK BIZ-UNK 0.75 10 10000 Number of alternatives (k) 2500 1 1000 KN-UNK BIZ-UNK Number of alternatives (k) 100 1000 10000 1 Number of alternatives (k) Figure 3 100 Number of alternatives (k) KN-UNK BIZ-UNK 250 10 10 100 1000 10000 Number of alternatives (k) Common unknown variance: Rows shows PCS and E[N ]/k vs. the number of alternatives k under the following configurations: SC (row 1); MDM (row 2); and RPI (row 3). For SC and MDM, P ⇤ = 0.9. For RPI, P ⇤ = 0.8. Sampling variances are common and unknown. E[N ]/k under BIZ initially increases with k, taking an average value near 800 at k = 10, and then declines for k 10 down to near 600 for the largest problems. In contrast, the average number of samples taken by KN is near 1400 at k = 10, and rises to near 2000 at k = 16, 384. Here again, BIZ’s ability to deliver a PCS that is closer to the target allows it to take fewer samples. Common unknown variance: Figure 3 shows the performance of KN and BIZ (Algorithm 2) on the configurations SC, MDM, and RPI, with unknown sampling variance. Although the vari- Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 29 ance is common across the alternatives, this fact is unknown to the algorithms. In this situation, we set n0 = 100 for both KN and BIZ, and set Bx = 1 for BIZ. Quantities are estimated using 10, 000 independent replications. Performance trends from the common known variance setting are apparent here as well: compared with KN, BIZ has a PCS that is closer to the target of P ⇤ = 0.9, and takes fewer samples. One di↵erence appears in the MDM configuration: the limiting value of E[N ]/k as k grows is n0 = 100, because each alternative must be sampled at least n0 times. These experimental results show that BIZ works well, even when the variance is unknown, at least in the settings investigated. Heterogeneous unknown variance: Figure 4 shows the performance of KN and BIZ (Algorithm 2 on problems with heterogeneous and unknown variance. The problem configurations considered are the following modifications of previously considered configurations. First, the Slippage Configuration with Increasing Variance (SC-INC) uses the same sampling means as the SC considered previously, µ1 = 2 x = (1 + x 1 ) 2, k 1 and µx = 0 for x 6= 1, and heterogeneous increasing sampling variances, with so that 2 k =2 DEC) is analogous, but uses 2 x 2 1. The Slippage Configuration with Decreasing Variance (SC- = (1 + k x ) 2, k 1 so 2 1 =2 2 k. Second, the Monotone Decreasing Means Configuration with Increasing Variance (MDM-INC) uses the same means as MDM, and the same variances as SC-INC. Similarly, the Monotone Decreasing Means Configuration with Decreasing Variance (MDM-DEC) uses the same means as MDM, and the same variances as SC-DEC. Third, Random Problem Instances with Heterogeneous Variances (RPI-HET) uses the same set of randomly generated means as RPI, and the same sampling variances as SC-INC and MDM-INC. In each configuration, performance quantities were estimated with 10, 000 independent replications. Again, BIZ performs well on these problem configurations, providing a PCS closer to P ⇤ and taking fewer samples than KN, providing evidence that BIZ’s performance is robust to both unknown and heterogeneous variance. Additional numerical experiments: Additional numerical experiments investigating the probability of good selection (Nelson and Banerjee 2001) for configurations outside the preference zone are presented in the appendix. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 30 2500 1500 PCS E[N]/k KN-UNK 2000 BIZ-UNK 1000 500 0 1 10 100 1000 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 10000 KN-UNK BIZ-UNK 1 Number of alternatives (k) 2500 1500 PCS E[N]/k KN-UNK 2000 BIZ-UNK 1000 500 0 1 10 100 1000 10000 PCS E[N]/k 200 100 0 10 100 1000 10000 PCS E[N]/k 200 100 0 10 100 1000 1 1000 10000 10 100 1000 10000 1 0.95 2000 PCS E[N]/k 100 Number of alternatives (k) KN-UNK 3000 BIZ-UNK 2500 1500 1000 0.9 0.85 0.8 500 0 KN-UNK BIZ-UNK 0.75 10 10000 KN-UNK BIZ-UNK Number of alternatives (k) 100 1000 Number of alternatives (k) Figure 4 10 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 10000 3500 1 1000 Number of alternatives (k) 300 1 100 KN-UNK BIZ-UNK 1 KN-UNK BIZ-UNK 400 10 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 Number of alternatives (k) 500 10000 Number of alternatives (k) 300 1 1000 KN-UNK BIZ-UNK 1 KN-UNK BIZ-UNK 400 100 1 0.98 0.96 0.94 0.92 0.9 0.88 0.86 Number of alternatives (k) 500 10 Number of alternatives (k) 10000 1 10 100 1000 10000 Number of alternatives (k) Heterogeneous unknown variance: Row shows PCS and E[N ]/k as a function of the number of alternatives k under the following configurations: SC-INC (row 1); SC-DEC (row 2); MDM-INC (row 3); MDM-DEC (row 4); and RPI-HET (row 5). For SC-INC, SC-DEC, MDM-INC, and MDM-DEC, P ⇤ = 0.9. For RPI-HET, P ⇤ = 0.8. Sampling variances are heterogeneous and unknown. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 31 6. Conclusion We have developed a new IZ procedure, called the Bayes-inspired IZ (BIZ) procedure. In continuous time, our lower bound on its worst-case preference-zone probability of correct selection is tight, and in discrete time, numerical experiments demonstrate that our lower bound is close to the procedure’s true worst-case probability of correct selection. This is the first sequential elimination procedure with tight worst-case preference-zone bounds on probability of correct selection for more than 2 alternatives. These theoretical results assume that the sampling variances are known, and have a particular integer multiple structure. In practice, sampling variances are unknown and do not possess an integer multiple structure, but numerical experiments suggest that the procedure’s performance is robust to violations of these assumptions. The tightness of this lower bound allows the procedure to take fewer samples than other IZ procedures, especially for problems with large numbers of alternatives. While the BIZ procedure takes only as many samples as is required to meet the desired lower bound on probability of correct selection, other procedures like the KN procedure of Kim and Nelson (2001) deliver a true probability of correct selection that is much larger than requested, and consequently take many more samples than is needed. Thus, having a tight lower bound improves efficiency and allows the BIZ procedure to select the best with fewer samples. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 32 Appendix A: Proofs The suppositions ur⇤ = maxx ux and u 2 PZ( ) imply ur⇤ > ux for all x 6= r⇤ . Proof of Lemma 1. This implies that the event X ⇤ = x0 is identical to the event R(x0 ) = r⇤ . Before computing the probability that R(x0 ) = r⇤ , we first compute the probability that R = r for a generic r. Consider any fixed permutation r of the integers {1, . . . , k }. By Bayes rule, the conditional probability Qu {R = r | YtA , (R(x))x2A / } is proportional to the likelihood of YtA and (R(x))x2A given R = r. If r 2 / R, this likelihood is 0 because r is inconsistent with the observed / (R(x))x2A / . If r 2 R, then r is consistent with the observed (R(x))x2A / , and this likelihood is the same as the likelihood of YtA given R = r. This likelihood is Qu {YtA 2 dy | R = r } = (2⇡ 2 ) 1 |A| 2 exp " 2 1 X = c(r, y) exp 2 1 X 2 ) 1 |A| 2 exp h 1 2 2 P fixed values for (r(x))x2A / , it follows that yx ur(x) x2A yx ur(x) x2A where c(r, y) = (2⇡ 2 ! 2 # dy dy, i 2 2 y + u x r(x) . Because r 2 R is a permutation of u with x2A P 2 x2A ur(x) is identical for all r 2 R. Furthermore, does not depend upon r, so c(r, y) is identical for all r 2 R. P 2 x2A yx This allows us to rewrite the likelihood as 1 X Qu {YtA 2 dy | R=r} / exp 2 ! yx ur(x) dy. x2A Now let r vary over R. Because the event {R(x0 ) = r⇤ } is the union of all the events {R = r} with r(x0 ) = r⇤ , and only those r 2 R have nonzero likelihood, we have Qu {R(x) = 1 | YtA , (R(x))x2A / }= X r2R:r(x0 )=r⇤ exp 1 X 2 x2A Ytx ur(x) ! X r2R exp 1 X 2 x2A ! Ytx ur(x) , which recovers the claimed expression. Proof of Lemma 2. Recall that Q = Qu , where u = [ , 0 . . . , 0] 2 PZ( ). Let R be as defined in Lemma 1 and let r⇤ = 1 2 arg maxx ux . Then, the expression from Lemma 1 provides an expression for Qu {X ⇤ = x0 | YtA , (R(x))x2A / }. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 On the event X ⇤ 2 A, and for r 2 R, the alternative r 1 P x2A Ytx ur(x) 33 = Yt,r 1 (r ) ⇤ (r⇤ ) which is best under permutation r, and ur since ur(x) = 0 for all x except = . The event X ⇤ 2 A is 1 (r ) ⇤ ⇤ measurable given (R(x))x2A / A. On this event, the expression / because X 2 A i↵ R(x) 6= r⇤ for all x 2 from Lemma 1 becomes ⇤ Qu {X ⇤ = x0 | YtA , (R(x))x2A / , X 2 A} = X exp r2R:r(x0 )=r⇤ X = exp r2R:r(x0 )=r⇤ = ax0 exp ✓ 2 ✓ ✓ Ytx0 Y 0 2 tx 2 ◆ Ytx0 ◆ X r2R ◆ X exp X ✓ Y 2 t,r X 1 (r ) ⇤ exp x2A r2R:r(x)=r⇤ ax exp x2A ✓ 2 ◆ ✓ ◆ 2 Ytx ◆ Ytx , where ax = | {r 2 R : r(x) = r⇤ } | is the number of elements of R in which a given alternative x is 1)! is constant for all x 2 A on the event X ⇤ 2 A, we cancel it to obtain the best. Since ax = (|A| ⇤ 0 ⇤ Qu {X = x | YtA , (R(x))x2A / , X 2 A} = exp ✓ 2 Ytx0 ◆ X exp x2A ✓ 2 ◆ Ytx . Finally, this expression does not depend upon (R(x))x2A / , and so ⇤ 0 ⇤ Qu {X = x | YtA , X 2 A} = exp Proof of Lemma 3. ✓ 2 Ytx0 ◆ X x2A exp ✓ 2 ◆ Ytx . Without loss of generality we assume that maxx ux = . If this is not the case, then we add the constant maxx0 ux0 to each ux , which does not change the value of Q u {X ⇤ = x | B }. Let x̃ 2 arg maxx0 2A yx0 . By the expression in Lemma 2, maxx2A Q {X ⇤ = x | B } = Q {X ⇤ = x̃ | B }. We will show that Q {X ⇤ = x̃ | B } Qu {X ⇤ = x̃ | B }, which is sufficient to show (5) because Qu {X ⇤ = x̃ | B } maxx2A Qu {X ⇤ = x | B }. For x 2 A, define fu (x) = Qu {X ⇤ = x | B } /Qu {X ⇤ = x̃ | B }. This quantity is well-defined because Qu {X ⇤ = x̃ | B } > 0. Then, Qu {X ⇤ = x̃ | B } = P Qu {X ⇤ = x̃ | B } 1 =P . ⇤ x2A Qu {X = x | B } x2A fu (x) P Similarly define f (x) = Q {X ⇤ = x̃ | B } /Q {X ⇤ = x | B } so that Q {X ⇤ = x̃ | B } = [ x2A f (x)] 1 . P Since [ x2A zx ] 1 is decreasing in each zx , to show that Qu {X ⇤ = x̃ | B } Q {X ⇤ = x̃ | B }, it is Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 34 enough to show that fu (x) f (x) for each x 2 A. This follows trivially for x = x̃, so we now consider x 6= x̃. Let x0 2 A with x0 6= x̃. We have, by Lemmas 2 and 1, where R and r⇤ are as defined in Lemma 1, P P 1 yx ur(x) 2 1 r2R:r(x̃)=r⇤ exp Px2A =P 1 0 f (x ) r2R:r(x0 )=r⇤ exp( 2 x2A yx ur(x) ) 1 fu (x0 ) exp exp( yx̃ 2 yx 0 ) 2 Multiplying through by the strictly positive quantity 2 4 X 1 X exp 2 r2R:r(x0 )=r⇤ x2A shows that the sign of this expression for [fu (x0 )] X exp 2 yx 0 + r2R:r(x̃)=r⇤ = X X j6=r⇤ r2R:r(x̃)=r⇤ ,r(x0 )=j = X 2 1 X 2 yx ur(x) x2A 0 4exp @ 2 0 exp @ X j6=r⇤ r2R:r(x̃)=r⇤ ,r(x0 )=j 0 exp @ ! yx 0 + !3 ✓ yx ur(x) 5 exp [f (x0 )] 1 X exp 1 2 is the same as the sign of yx̃ + r2R:r(x0 )=r⇤ u r⇤ 2 yx̃ + uj 1 X X 1 2 x2A\{x̃,x0 } 2 yx̃ + u r⇤ 2 yx 0 + uj 2 2 yx̃ + 1 2 X y + 2 x̃ X 1 2 x2A\{x̃,x0 } 1 1 ! yx ur(x) A x2A\{x̃,x0 } y 0+ 2 x yx ur(x) x2A yx 0 + 2 yx 0 2 ◆ 13 yx ur(x) A5 h ⇣u ⌘ j yx ur(x) A exp 2 yx0 exp ⇣u j y 2 x̃ ⌘i . (11) The second line uses that the values of r(x) for x 2 A \ {x̃, x0 } are the same between the two sums P r2R:r(x̃)=r⇤ ,r(x0 )=j and P r2R:r(x0 )=r⇤ ,r(x̃)=j . The third line uses that ur⇤ = . Consider the sign of the last term, exp(uj yx0 / u 2 PZ( ) imply uj 0. Furthermore, yx̃ implies exp(yx0 uj / 2 ) exp(Ytx̃ uj / 2 ) 2 ) exp(uj yx̃ / 2 ). Together ur⇤ = , j 6= r⇤ and yx0 then implies exp(uj yx̃ / 2 ) exp(uj yx0 / 2 ), which 0. Thus, the sign of (11) is nonnegative, which shows that fu (x) f (x) for each x, which shows (5). The proof of (6) follows a similar argument, but with x̃ 2 arg minx2A yx . In this case, fu (x) for each x 2 A because yx̃ yx0 implies that exp(yx0 uj / positive. 2 ) exp(yx̃ uj / 2 f (x) ) 0 and (11) is non- Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 Proof of Lemma 4. 35 We will show recursively that ⌧n < 1 a.s. for each n = 0, 1, . . . , k statement holds for ⌧0 = 0. Now suppose that ⌧n < 1 a.s. for some n k 1. The 2, and we will show that ⌧n+1 < 1 a.s. 2 Define a random variable a = 1 log((k Pn )). If n = 0 then Pn = P ⇤ 2 (0, 1), 1)Pn )/(1 implying a is finite. If n > 0, then P ⇤ Pn P ⇤ /(1 c)n P ⇤ /(1 c)k 2 P ⇤ /(P ⇤ )(k 2)/(k 1) < 1, implying Pn 2 (0, 1) and a is finite. Fix a deterministic time t 0, let ⌧ 0 = ⌧n + t, and let x̃ be any F⌧ 0 -measurable random variable that is almost surely in An . Consider the event that ⌧n < 1 and Y⌧ 0 +1,x̃ x 2 An \ {x̃}. On this event, q⌧ 0 +1,x (An )/q⌧ 0 +1,x̃ (An ) = exp( x 2 An \ {x̃} and 2 q⌧ 0 +1,x̃ (An ) = 41 + X x2An \{x̃} 3 2 Y⌧ 0 +1,x a for each Y⌧ 0 +1,x̃ )) exp( a / (Y⌧ 0 +1,x 2 ) for 1 q⌧ 0 +1,x (An )/q⌧ 0 +1,x̃ (An )5 ⇥ 1 + (k 1) exp( a / 2 ) ⇤ 1 = Pn . The value of a was chosen to make this last equality with Pn hold. Thus, on the event considered, q⌧ 0 +1,x̃ Pn and ⌧n+1 ⌧ 0 + 1. We now define x̃ 2 arg maxx2An Y⌧ 0 ,x , which is F⌧ 0 -measurable and is almost surely in An . The previous discussion implies Pµ, {⌧n+1 ⌧ 0 + 1 | F⌧ 0 , ⌧n+1 > ⌧ 0 } Pµ, {Y⌧ 0 +1,x̃ a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 } . Y⌧ 0 +1,x This implies that Pµ, {⌧n+1 ⌧ 0 + 1 | F⌧ 0 , ⌧n+1 > ⌧ 0 } Pµ, {Y⌧ 0 +1,x̃ Y⌧ 0 ,x̃ , Y⌧ 0 ,x̃ Y⌧ 0 +1,x a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 } Pµ, {Y⌧ 0 +1,x̃ Y⌧ 0 ,x̃ , Y⌧ 0 ,x Y⌧ 0 +1,x = Pµ, {Y⌧ 0 +1,x̃ Y⌧ 0 ,x̃ | F⌧ 0 } Y a 8x 2 An \ {x̃} | F⌧ 0 , ⌧n+1 > ⌧ 0 } where the second line follows from the fact Y⌧ 0 ,x̃ x2An \{x̃} Pµ, {Y⌧ 0 ,x Y⌧ 0 +1,x a | F⌧ 0 } . Y⌧ 0 ,x and the third line follows from the inde- pendence of increments of Ysx under Pµ, . The probability Pµ, {Y⌧ 0 +1,x̃ variable, Y⌧ 0 +1,x̃ Y⌧ 0 ,x̃ | F⌧ 0 } is the probability of a conditionally N (µx̃ , Y⌧ 0 ,x̃ , exceeding 0. This probability is (µx̃ / 2 2 ) random ), which is bounded below by Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 36 (minx µx / Pµ, {Y⌧ 0 ,x Yt 2 Y⌧ 0 +1,x Yt,x 1,x maxx0 µx0 )/ ). Here, is the normal cumulative distribution function. Similarly, the probability a | F⌧ 0 } is the probability of a conditionally N ( µx a, exceeding 0. This probability is ( (a+µx )/ 2 2 a, 2 ) random variable, ), and is bounded below by ( (a+ ). Thus, replacing ⌧ 0 with ⌧n + t, Pµ, {⌧n+1 ⌧n + t + 1 | F⌧n +t , ⌧n+1 > ⌧n + t} 2 (min µx / x ) h ( (a + max µx )/ x 2 ) ik 1 . Let ✏ be the quantity on the right-hand side of this inequality. ✏ < 1 is an F⌧n -measurable random variable that does not depend on t. By repeated application of this inequality, we have for any finite t that Pµ, {⌧n+1 > ⌧n + t | F⌧n } (1 ✏)t , which vanishes in the limit as t ! 1. This implies that Pµ, {⌧n+1 = 1} = 0. Proof of Lemma 5. We show the claim using induction. It is trivially true for ⌧0 = 0. Now suppose the claim is true for some n, and we show that it holds for n + 1. To show that ⌧n+1 is a stopping time, it suffices to show that {⌧n+1 t} is Ft measurable for each t 0. Let t 0, and write {⌧n+1 t} = {⌧n t} \ {qsx (An ) 2 / (c, Pn ) for some s 2 [⌧n , t] \ T } = [ A✓{1,...,k} {⌧n t} \ {An = A} \ {qsx (A) 2 / (c, Pn ) for some s 2 [⌧n , t] \ T } The last event in this expression {qsx (A) 2 / (c, Pn ) for some s 2 [⌧n , t] \ T } can be rewritten as S s2Q {s 2 [⌧n , t] \ T} \ {qsx (A) 2 / (c, Pn )} where Q denotes the rational numbers. For the case T = Z+ , this follows from T ⇢ Q. For the case T = R+ , it follows from the fact that (c, Pn ) is an open set, and from the continuity of the sample paths of s 7! qsx (A). (This sample-path continuity follows in turn from the continuity of qsx (A) considered as a function of Ys , and the continuity of the sample paths of Ys .) It is convenient to rewrite and summarize this as {qsx (A) 2 / (c, Pn ) for some s 2 [⌧n , t] \ T } = [ s2Q\T\[0,t] {⌧n s} \ {qsx (A) 2 / (c, Pn )} Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 37 Now consider the term {qsx (A) 2 / (c, Pn )}. This can be rewritten as {qsx (A) 2 / (c, Pn )} = \ {p p2Q Pn } [ {qsx (A) 2 / (c, p)} . Putting this all together, {⌧n+1 t} = [ \ A✓{1,...,k} p2Q s2Q\T\[0,t] [ = \ A✓{1,...,k} p2Q s2Q\T\[0,t] {⌧n t} \ {An = A} \ {⌧n s} \ {p Pn } [ {qsx (A) 2 / (c, p)} , {⌧n t, An = A} \ {⌧n s} \ {⌧n t, p Pn } [ {qsx (A) 2 / (c, p)} . In this expression, {⌧n t, An = A} is Ft -measurable because An is F⌧n measurable; {⌧n t, p Pn } is Ft -measurable because Pn is F⌧n measurable; {⌧n s} is Ft -measurable because ⌧n is a stopping time and s t; and {qsx (A) 2 / (c, p)} is Ft -measurable because qsx (A) is Fs -measurable and s t. Since countable unions and intersections of Ft -measurable events are Ft -measurable, {⌧n+1 t} is Ft -measurable. Part (a): Consider the event {M Proof of Lemma 6. M =k k 1}. On this event, we show that Pk 2 . (12) 1. To show this, it is sufficient to show max q⌧k 1 ,x x2Ak 2 We know ⌧k 1 (Ak 2 ) < 1 almost surely by Lemma 4, and the definition of ⌧k 1 implies that either (12) holds or the following relation holds, min q⌧k x2Ak 2 The number of elements in Ak 2 1 ,x (Ak 2 ) c. (13) is exactly 2. Call them x(1) and x(2) , where Y⌧k 1 ,x (1) Y⌧k 1 ,x (2) . By Lemma 2, for x 2 x(1) , x(2) , q⌧ k and q⌧k 1 ,x (1) q⌧ k 1 ,x (2) 1 ,x (Ak 2 ) = exp( exp( 2 Y⌧k Y⌧k 1 ,x ) (1) ) + exp( 2 Y⌧ 1 ,x k 2 1 ,x (2) ) ◆ . , . The expression (12) holds i↵ exp ✓ 2 Y⌧k 1 ,x (1) ◆ Pk 2 1 Pk exp 2 ✓ 2 Y⌧k 1 ,x (2) (14) Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 38 Similarly, the expression (13) holds i↵ ✓ ◆ exp Y (1) 2 ⌧k 1 ,x On the event {M Pk c c P⇤ (1 c)k 2 = (1 c) P⇤ (1 c)k 1 where the first inequality is due to Pn = Pn 1 /(1 1, . . . , k 2 on {M k exp ✓ Y 2 ⌧k 1 ,x (2) ◆ . (15) 1}, k 2 1 (1 c) P⇤ (P ⇤ )(k minx2An 1 1)/(k 1) 2 c, q⌧n,x (An 1 )) Pn /(1 1}, and the second inequality is due to c 1 Pk 2 1 Pk =1 (P ⇤ )1/(k 1) c) for n = . This implies 1 c 1 c = , 1 (1 c) c implying that if (15) holds, so does (14). The fact that at least one of (15) and (14) holds implies that (14) holds, which implies that (12) holds. This shows part (a). Part (b): Let Y ⇤ = maxx2AM 1 Y⌧M ,x . We claim that, for n = M, M + 1, . . . , k 1, the following statements hold: (i) ⌧n = ⌧M . (ii) maxx2An Y⌧n ,x = Y ⇤ . (iii) maxx2An 1 q⌧n ,x (An 1 ) Pn 1 . We show this by induction on n. We first consider the base case, n = M . In this case, (i) follows trivially. (ii) follows because AM = AM 1 \ {ZM }, ZM 2 arg minx2AM 1 Y⌧M ,x and |AM 1| =k (M 1) 2 by part (a). (iii) follows from the definition of M . We now show the induction step. Suppose that (i), (ii), and (iii) hold for some n 2 [M, k 2]. We show they also hold for n + 1. We have, by (ii) for n, " P # 1 exp( Y ) exp( 2 Y ⇤ ) exp( 2 Y ⇤ ) ⌧ ,x 2 n max q⌧n ,x (An ) = P =P · P x2An . (16) x2An exp( Y ) exp( Y ) exp( Y ⌧ ,x ⌧ ,x ⌧n ,x ) 2 2 2 n n x2An x2An 1 x2An 1 P The first term of (16) satisfies exp( 2 Y ⇤ )/ x2An 1 exp( 2 Y⌧n ,x ) = maxx2An 1 q⌧n ,x (An 1 ) Pn 1 , where we have used (iii) for n. We rewrite the ratio in the second term of (16) as P exp( 2 Y⌧n ,x ) exp( 2 Y⌧n ,Zn ) P x2An =1 P = 1 q⌧n ,Zn (An 1 ) = 1 min q⌧n ,x (An 1 ), x2An 1 x2An 1 exp( 2 Y⌧n ,x ) x2An 1 exp( 2 Y⌧n ,x ) Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 where we use that An 1 39 \ An = {Zn } in the first equality, the definition of qt,x (·) in the second equality, and that Zn 2 arg minx2An q⌧n ,x (An 1 ) in the third equality. Combining these two, (16) 1 becomes max q⌧n ,x (An ) Pn 1 [1 x2An 1 min q⌧n ,x (An 1 )] = Pn . x2An 1 This and the definition of ⌧n+1 implies that ⌧n+1 = ⌧n , implying in turn that (i) and (iii) are met for n + 1. Finally, Zn+1 2 arg minx2An Y⌧n+1 ,x = arg minx2An Y⌧M ,x , and An+1 = An \ {Zn+1 }, so maxx2An+1 Y⌧n+1 ,x = maxx2An+1 Y⌧M ,x = maxx2An Y⌧M ,x = Y ⇤ and (iii) is met for n + 1. This shows that (i),(ii), and (iii) are met for n = M, . . . , k in Ak 1 , and (i),(ii) for n = k maxx2AM 1 1 implies that this element must have Y⌧k Y⌧M ,x , so x̂ 2 arg maxx2AM Part (c): To show that ⌧M > ⌧M PM 1. 1. Finally, x̂ is the single element = Y⌧M ,x̂ = Y ⇤ = 1 ,x̂ Y⌧M ,x . This shows part (b). 1 1, it is sufficient to show that maxx2AM 1 q⌧ M 1 ,x arg maxx2AM Y⌧M 1 ,x max q⌧M 1 ,x 2 x2AM 2 =P exp( x2AM = q⌧ M P < 1 ,x̃ (AM 2) Y⌧M 1 ,x̃ ) exp( 2 Y⌧M 2 1) 1 1 = AM 1 exp( 2 Y⌧M 1 ,x ) x2AM 2 exp( 2 Y⌧M 1 ,x ) q⌧ M 1 ,ZM 1 1 ,x q⌧ M x2AM used that 1 = q⌧ M 2 (AM =1 (AM 2) ) 1 ,x̃ =P 1 ,ZM 2 (AM 1 2) exp( Y⌧M 1 ,x̃ ) exp( 2 Y⌧M 1 x2AM (AM 2) \ {Z M 1} exp 2 Y ⌧M P ⇣ x2AM =1 2 2 = q⌧M 1 ,x̃ 1 ,x ) (AM P ·P x2AM 1 x2AM 2 1 )PM exp( 2 Y⌧M 1 ,x ) exp( 2 Y⌧M 1 ,x ) 2 /PM 1. in going from the second to the third line to write 1 ,ZM 1 ⌘ exp( 2 Y⌧M 1 ,x ) minx2AM 2 q⌧M 1 ,x =1 q⌧ M (AM 2) 1 ,ZM = PM 1 (AM 2 /PM 1 2 ). 1 ,x̃ (AM 1) = maxx2AM 2 q⌧M 1 ,x (AM PM 2 /PM 1 where the inequality follows from the definition of M . 2) < PM PM 2 2 /PM = PM 1 We have also to write the last equality on the third line. From this we have, q⌧ M 2. Let x̃ 2 , so where we have used AM that 1) We show this by considering two cases. In the first case, suppose M = 1. Then the result follows because maxx2A0 q0x = 1/k < P ⇤ = P0 . In the second case, suppose M P (AM 1, Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 40 Finally, because an alternative with minimal Y⌧M we have arg maxx2AM maxx2AM 1 q⌧M 1 ,x (AM 2 Y⌧M 1) 1 ,x = q⌧ M = arg maxx2AM 1 ,x̃ (AM 1) 1 1 ,x Y⌧M < PM 1. is eliminated in going from AM 1 ,x 2 to AM 1, , and x̃ is an element of this set. Thus, This shows part (c). Proof of Lemma 7. Because Ytx is almost surely finite at any fixed t, qtx (A0 ) > 0 almost surely. Thus, when c = 0, minx2An qtx (An ) > 0 = c. This implies that ⌧1 = inf {t 2 T : maxx2A0 qtx (A0 ) P0 } and M = 1. By Lemma 6, ⌧ = ⌧1 . Recalling P0 = P ⇤ , A0 = {1, . . . , k }, and the definition of qtx (A), allows us to rewrite ⌧ as claimed. Proof of Lemma 8. The invariance of PCS(u, ) to permutations of u follows from the invari- ance of the BIZ procedure to the labeling of the alternatives. Let a 2 R, and we now show that PCS(u, ) = PCS(u + [a, . . . , a], ), so PCS(u, ) is also invariant to translations of u. We work under the probability measure Pµ, with µ = u + [a, . . . , a]. Let Wtx = (Ytx tux ta))/ so that {Wtx : t 2 R+ , x 2 {1, . . . , k }}, is a k-dimensional standard Brownian motion. We then write qtx (A) in terms of Wtx . From (2), and Ytx = Wtx + tux + ta, qtx (A) = exp = exp ✓ ✓ 2 ( Wtx + tux + ta) ( Wtx + tux ) 2 ◆ ◆ X x0 2A X x0 2A exp exp ✓ ✓ 2 ( Wtx + tux + ta) ◆ ( W + tu ) , tx x 2 where we have divided the numerator and denominator by exp( ta/ 2 ◆ ). Thus, the path of the stochastic process (qtx (A) : t 2 R+ , A ✓ {1, . . . , k }) can be written as a function of the path of a k-dimensional standard Brownian motion, and this function does not depend upon a. Furthermore, since the event of correct selection can be written entirely in terms of the paths of qtx (A), the probability of correct selection PCS(u + [a, . . . , a], ) also does not depend upon a. Now we show that Qu {CS} = PCS(u, ). First, rewrite the probability of CS under Qu as Qu {CS} = X r Qu {R = r} Qu {CS | R = r} . Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 41 Conditioning on R = r under Qu is the same as conditioning on ✓ = ũ where ũx = ur(x) for each x. Thus, Qu {CS | R = r} = Qu {CS | ✓ = ũ} = Pũ, {CS } = PCS(ũ, ). Since PCS(·, ) is invariant to permutations of the alternatives, and ũ is a permutation of u, PCS(ũ, ) = PCS(u, ). This implies Qu {CS} = X r Qu {R = r} PCS(u, ) = X r ! Qu {R = r} PCS(u, ) = PCS(u, ), because Qu {R = r} sums over r to 1. Proof of Lemma 9. Fix u 2 PZ( ). First, since ⌧k 1 < 1 almost surely under any Pv, with v 2 Rk by Lemma 4, and Qu is a mixture of such Pv, , Qu {CS} = Qu {x̂ = X ⇤ , ⌧k 1 < 1} = Qu {x̂ = X ⇤ }. Thus we may work with the probability of {x̂ = X ⇤ } rather than the probability of CS. To simplify notation, let Gn be the sigma-algebra generated by F⌧n and (R(x))x2A / n 1 . Let Cn be the event {X ⇤ 2 An 1 , M = n} and Dn be the event {X ⇤ 2 An 1 , M > n}. Both events are measurable with respect to Gn . To show the lemma, it is enough to show Qu {x̂ = X ⇤ | Gn , Cn } Pn 1 for n = 1, . . . , k 1, (17) Qu {x̂ = X ⇤ | Gn , Dn } Pn 1 for n = 1, . . . , k 2, (18) with equality if u = [ , 0, . . . , 0] and T = R+ . We first show (17), and then show (18). Proof of (17). On the event Cn ✓ {M = n}, Lemma 6 implies x̂ 2 arg maxx2An Lemma 1, this argmax set is identical to arg maxx2An 1 1 Y⌧n ,x . By Qu {X ⇤ = x | Gn }. This implies that, Qu {x̂ = X ⇤ | Gn , Cn } = max Qu {X ⇤ = x | Gn , Cn } x2An 1 max Q {X ⇤ = x | Gn , Cn } = max q⌧n ,x (An 1 ) x2An 1 x2An 1 Pn 1 . On the second line, the first inequality is due to Lemma 3, the equality is due to the definition of qtx (A), and the second inequality is due to the definition of M and Cn ✓ {M = n}. If u = [ , 0, . . . , 0], then Qu = Q and the first inequality is an equality. If T = R+ , the second inequality is an equality due to the continuity of the paths of Brownian motion and ⌧M 1 < ⌧M , as shown in Lemma 6. Thus, if u = [ , 0, . . . , 0] and T = R+ , then Qu {x̂ = X ⇤ | Gn , Cn } = Pn 1 . Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 42 Proof of (18). We now show (18). The proof is by backward induction on n. We first consider n=k 2. Lemma 6 part (a) implies that M = k (17) implies (18) holds when n = k Now fix some n < k 1 almost surely on the event M > k 2. Then, 2, and with equality if u = [ , 0, . . . , 0] and T = R+ . 2. Our induction hypothesis is that (18) holds for n + 1, and that it holds with equality if u = [ , 0, . . . , 0] and T = R+ . We will show this is then also true for n. We have, Qu {x̂ = X ⇤ | Gn , Dn } = Qu {Zn 6= X ⇤ | Gn , Dn } Qu {x̂ = X ⇤ | Gn , Dn , Zn 6= X ⇤ } . (19) We rewrite the first term in (19) as Qu {Zn 6= X ⇤ | Gn , Dn } = 1 =1 1 =1 Q u {Z n = X ⇤ | G n , D n } min Qu {x = X ⇤ | Gn , Dn } x2An 1 min Q {x = X ⇤ | Gn , Dn } x2An 1 min q⌧n ,x (An 1 ) = Pn 1 /Pn . x2An 1 The second line follows from Zn 2 arg minx2An 1 Qu {x = X ⇤ | Gn }. The third line follows from Lemma 3, and is equality if u = [ , 0, . . . , 0]. The fourth and last line follows from the definition of ⌧n and that we condition on Dn ✓ {M > n}. We now consider the second term of (19). Define the event E = Dn \ {Zn 6= X ⇤ } = {X ⇤ 2 / An 1 , Zn 6= X ⇤ , M > n} = {X ⇤ 2 An , M n + 1} = Cn+1 [ Dn+1 . The second term of (19), Qu {x̂ = X ⇤ | Gn , Dn , Zn 6= X ⇤ }, can be rewritten, Qu {x̂ = X ⇤ | Gn , E } = EQu [Qu {x̂ = X ⇤ | Gn+1 , E } | Gn , E] EQu [Pn | Gn , E] = Pn , where EQu indicates the expectation taken with respect to Qu . In the first equality we have used the tower property of conditional expectation and Gn ⇢ Gn+1 . In the inequality, we have used that (17) implies Qu {x̂ = X ⇤ | Gn+1 , Cn+1 } Qu {x̂ = X ⇤ | Gn+1 , Dn+1 } Pn and the induction hypothesis implies Pn , which together imply Qu {x̂ = X ⇤ | Gn+1 , Cn+1 [ Dn+1 } inequality is equality if T = R+ and u = [ , 0, . . . , 0]. Pn . This Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 43 Combining the two terms of (19), we obtain Pn 1 Pn = Pn , Pn Qu {x̂ = X ⇤ | Gn , Dn } with equality if u = [ , 0, . . . , 0] and T = R+ . Thus, the induction statement holds. Appendix B: Additional Numerical Results In Figure 5 we test the performance of BIZ on problem configurations outside the preference zone. Recall that a procedure satisfying the indi↵erence-zone guarantee requires that the lower bound P ⇤ on PCS only be met for problem configurations in the preference zone PZ( ) = µ 2 Rk : µ[k] µ[k 1] , where the best alternative is better than the second best alternative by at least . For studying problem configurations outside the preference zone, Nelson and Banerjee (2001) suggests that performance be measured by the probability of good selection (PGS), PGS(µ, ) = Pµ, n x̂ max µx x o , ⌧ <1 . For configurations in the preference zone, PGS is equal to PCS. But for configurations outside the preference zone, the two quantities di↵er. A natural generalization of the indi↵erence-zone guarantee would be a lower bound on PGS over all configurations, i.e., a probability of good selection guarantee of the form, PGS(µ, ) P⇤ for all µ 2 Rk+ . A procedure that satisfies this PGS guarantee also satisfies the IZ guarantee, but a procedure satisfying the IZ guarantee does not necessarily satisfy the PGS guarantee. The numerical results in Figure 5 suggest that the BIZ procedure may satisfy the PGS guarantee, at least in the common known variance case. However, this has not been confirmed theoretically, and doing so is left for future work. Acknowledgments The author was supported by AFOSR YIP FA9550-11-1-0083 and NSF CAREER CMMI-1254298. The author would like to thank Rolf Waeber and Shane Henderson for helpful discussions, and the associate editor and several anonymous referees for helpful comments. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 PGS 44 0.99 0.98 0.97 0.96 0.95 0.94 0.93 0.92 0.91 0.9 0.89 k=3 k=10 k=100 0 0.5 1 1.5 2 (mu(1) - mu(2)) / Delta Figure 5 The probability of good selection (PGS) plotted as a function of (µ1 µ2 )/ for BIZ procedure with common known-variance for problem configurations with k = 3, 10, and 100 alternatives. All configurations considered had µ = [ , µ2 , 0, . . . , 0], = 0.5, P ⇤ = 0.9, and = 10. When (µ1 have a slippage configuration with parameter . When this quantity is µ2 )/ = 1, we 1, the configuration is in the preference zone, and when it is < 1, it is in the indi↵erence zone. References Andradóttir, S., S.H. Kim. 2010. Fully sequential procedures for comparing constrained systems via simulation. Naval Research Logistics 57(5) 403–421. Bechhofer, R.E. 1954. A single-sample multiple decision procedure for ranking means of normal populations with known variances. The Annals of Mathematical Statistics 25(1) 16–39. Bechhofer, R.E., J. Kiefer, M. Sobel. 1968. Sequential Identification and Ranking Procedures. University of Chicago Press, Chicago. Bechhofer, R.E., T.J. Santner, D.M. Goldsman. 1995. Design and Analysis of Experiments for Statistical Selection, Screening and Multiple Comparisons. J.Wiley & Sons, New York. Berger, James O. 1985. Statistical decision theory and Bayesian analysis. 2nd ed. Springer-Verlag, New York. Branke, J., S.E. Chick, C. Schmidt. 2007. Selecting a selection procedure. Management Science 53(12) 1916–1932. Chick, S.E., J. Branke, C. Schmidt. 2010. Sequential sampling to myopically maximize the expected value of information. INFORMS Journal on Computing 22(1) 71–80. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 45 Chick, S.E., K. Inoue. 2001. New two-stage and sequential procedures for selecting the best simulated system. Operations Research 49(5) 732–743. Dieker, AB, Seong-Hee Kim. 2012. Selecting the best by comparing simulated systems in a group of three when variances are known and unequal. Proceedings of the 2012 Winter Simulation Conference. IEEE. Fabian, V. 1974. Note on Anderson’s sequential procedures with triangular boundary. The Annals of Statistics 2(1) 170–176. Frazier, P., W. B. Powell, S. Dayanik. 2008. A knowledge gradient policy for sequential information collection. SIAM Journal on Control and Optimization 47(5) 2410–2439. Frazier, P., W. B. Powell, S. Dayanik. 2009. The knowledge gradient policy for correlated normal beliefs. INFORMS Journal on Computing 21(4) 599–613. Frazier, P.I., W.B. Powell. 2008. The knowledge-gradient stopping rule for ranking and selection. S.J. Mason, R.R. Hill, L. Mönch, O. Rose, T. Je↵erson, J.W. Fowler, eds., Proceedings of the 2008 Winter Simulation Conference. Institute of Electrical and Electronics Engineers, Inc., Piscataway, New Jersey, 305–312. Goldsman, D., S.H. Kim, W.S. Marshall, B.L. Nelson. 2002. Ranking and selection for steady-state simulation: Procedures and perspectives. INFORMS Journal on Computing 14(1) 2–19. Gupta, S.S., K.J. Miescke. 1996. Bayesian look ahead one-stage sampling allocations for selection of the best population. Journal of Statistical Planning and Inference 54(2) 229–244. Hartmann. 1991. An improvement on Paulson’s procedure for selecting the population with the largest mean from k normal populations with a common unknown variance. Sequential Analysis 10(1-2) 1–16. Hartmann, M. 1988. An improvement on Paulson’s sequential ranking procedure. Sequential Analysis 7(4) 363–372. Hong, J. 2006. Fully sequential indi↵erence-zone selection procedures with variance-dependent sampling. Naval Research Logistics 53(5) 464–476. Kim, Seong-Hee, AB Dieker. 2011. Selecting the best by comparing simulated systems in a group of three. Proceedings of the 2011 Simulation Conference. IEEE, 3987–3997. Kim, Seong-Hee, Barry L. Nelson. 2001. A fully sequential procedure for indi↵erence-zone selection in simulation. ACM Trans. Model. Comput. Simul. 11(3) 251–273. Frazier: IZ R&S with Tight Bounds on Probability of Correct Selection Article submitted to Operations Research; manuscript no. OPRE-2011-09-480 46 Kim, S.H., B.L. Nelson. 2006. Selecting the best system. S.G. Henderson, B.L. Nelson, eds., Handbook in Operations Research and Management Science: Simulation. Elsevier, Amsterdam, 501–534. Kim, S.H., B.L. Nelson. 2007. Recent advances in ranking and selection. Proceedings of the 39th conference on Winter simulation: 40 years! The best is yet to come. IEEE Press, Piscataway NJ, 162–172. Malone, Gwendolyn J, Seong-hee Kim, David Goldsman, Demet Batur. 2005. Performance of Variance Updating Ranking and Selection Procedures. Proceedings of the 2005 Winter Simulation Conference. IEEE. Nelson, B., J. Swann, D. Goldsman, W. Song. 2001. Simple procedures for selecting the best simulated system when the number of alternatives is large. Operations Research 49(6) 950–963. Nelson, B.L., S. Banerjee. 2001. Selecting a good system: Procedures and inference. IIE Transactions 33(3) 149–166. Paulson, E. 1964. A sequential procedure for selecting the population with the largest mean from k normal populations. The Annals of Mathematical Statistics 35(1) 174–180. Paulson, E. 1994. Sequential procedures for selecting the best one of k Koopman-Darmois populations. Sequential Analysis 13(3). Rinott, Y. 1978. On two-stage selection procedures and related probability-inequalities. Communications in Statistics-Theory and Methods 7(8) 799–811. Swisher, J.R., S.H. Jacobson, E. Yücesan. 2003. Discrete-event simulation optimization using ranking, selection, and multiple comparison procedures: A survey. ACM Transactions on Modeling and Computer Simulation 13(2) 134–154. Wang, Huizhu, Seong-Hee Kim. 2012. On the conservativeness of fully sequential indi↵erence-zone procedures. Submitted for publication .
© Copyright 2026 Paperzz