Discriminating Methods: Nonnested Tests for Strategic Choice Models Kevin A. Clarke (585) 275-5217 [email protected] † Curt S. Signorino (585) 273-4760 [email protected] Department of Political Science University of Rochester Harkness Hall Rochester, NY 14627-0146 (585) 271-1616 (F) Abstract We consider a researcher attempting to choose, on the available data, between a strategic choice model and various nonstrategic choice models. Such a researcher is faced with choosing between rival models that are nonnested in terms of their functional forms. We discuss both a parametric and distribution-free procedure for making this choice, and demonstrate through a monte carlo simulation that discrimination is possible. The results of the simulation also allow us to compare the relative power of the two tests. May 4, 2003 † An earlier version of this paper was presented at the annual meeting of the International Studies Association in Portland, OR (2003); we thank the participants for their comments. Support from the National Science Foundation (Grant SES-0213771) is gratefully acknowledged. 1 Introduction The empirical study of political science has, in recent years, undergone something akin to a sea-change. Where it was once common for political scientists to employ a linear functional form regardless of the theory being tested, we now see a new attention being paid to the connection between theory and model. One result of this attention has been the realization that a disconnect between theory and model results in serious misspecification and parameter bias (Signorino and Yilmaz 2000). Now that researchers are encouraged to derive a functional form directly from the theoretical model under consideration, new challenges have arisen. Although Signorino (1999) has demonstrated that traditional specifications of statistical models are generally inconsistent with strategic theories of political science, no rigorous framework has emerged for comparing strategic models against one another, or against nonstrategic models. While it is clear that strategic specifications provide different answers from traditional specifications, it not yet clear that these strategic specifications are, in fact, superior. The question at hand is whether or not strategic specifications are “closer” than nonstrategic models to the data generating processes at work in the field of international relations. Compounding the difficulty of the issue is the fact that not all theory is detailed enough to allow the derivation of a functional form suitable for testing. Achen (1982), for example, argues that good social science theory is functionally non-specific. An empirical researcher faced with choosing a specification under these conditions needs guidance in choosing between the many functional forms, some strategic and some not, that may be used to model most political phenomena. In either case, empirical researchers need tools that allow models with different functional forms to be compared. Models with different functional forms, however, are generally nonnested (neither model is a special case of the other model). Discriminating between nonnested models requires specialized tests that are rarely used in political science research. Clarke (2001) introduced the issue of nonnested testing to political science, and Clarke (2003b) introduced a simple distribution-free test for nonnested model discrimination. These articles, however, consider only models that are nonnested in terms of their covariates. Testing models that are nonnested in terms of their functional forms is a natural extension of this line of research. This paper, in essence, represents a first step in combining the research programmes of its authors. Our hope is that future scholars will be able to develop statistical models that are truly consistent with their theories and be able to test those theories against one another. The attainment of these goals will allow the scientific study of international relations to move significantly beyond its current state. The paper proceeds as follows. In Section 2, we detail a discrete choice problem that may be modelled in a number of different ways. We demonstrate that the assumptions made by the researcher regarding the decision-making process determine the correct functional form. In Section 3, we review both a parametric and distribution-free discrimination test for nonnested models. Finally, in Section 4, we perform a suite of monte carlo simulations to demonstrate that discriminating between alternative choice models characterized by nonnested functional forms is feasible. The simulations also allow us to compare the relative power of the two tests. 2 Alternative Choice Models Consider a researcher who wishes to model the decision-making process of two individuals choosing between three unordered alternatives. A common approach to this problem is to associate choice with utility and assume that individual i receives utility Uij when alternative j is chosen. There exist, however, at least three common “random utility” specifications the researcher could use to model this particular choice problem. As a way of fixing ideas, let us think through a simple example in which a married couple is attempting to choose between three different modes of transportation for getting to work.1 The options include riding their bikes, taking a bus, or taking a taxi. A simple, but often overlooked, point is that how this example should be modelled depends on how the couple in question makes their decision. For instance, the researcher might assume that the couple makes their choice jointly. That is, the couple decides, as a unit, which transportation option provides them the greatest utility. The researcher should therefore model the situation as two individuals maximizing their joint utility and choosing option ∗ ∗ > Uik , ∀ k = j. j if Uij Choosing a mode of transportation as a unit is not, however, the only method two individuals could employ to make this decision. Another possibility is that the researcher could assume that the choice occurs sequentially. Individual 1 chooses between option j and not option j, leaving individual 2 to choose between options k and l. Returning to our example, the wife may choose between riding bikes or not based on her expected utility for these options, but not take into account her husband’s preferences regarding the bus versus a taxi. If the wife decides that the couple will not ride their bikes to work, her husband then may choose between taking a bus or taking a taxi based on his expected utility. Note the censoring we may observe given this arrangement. If the wife chooses bike riding, we do not learn anything about the choice the husband might have made. A third possibility is that the researcher might assume that the choice occurs not only sequentially, but strategically. That is, individual 1 chooses between option j and not option j, based not on an expected utility calculation as in the example above, but conditioning on his or her belief about individual 2’s 1 We will consider a more complex political example in the next section. 2 subsequent choice between k and l.2 Continuing our running example, the wife decides whether or not the couple should ride their bikes, but makes this choice taking into consideration whether she believes her husband will choose a taxi over a bus or vice versa. The husband then makes his choice between the bus or a taxi based on a straightforward expected utility calculation. This decisionmaking process is quite different from the last example where the wife does not take into account what she thinks her husband will do. As in the last example, if the wife chooses bike riding, we do not observe the husband’s choice. Were we to write down a statistical model for each of these situations, the three different assumptions made by the researcher regarding the couple’s decision-making process would lead to three different functional forms: the independent multinomial probit or logit, a nonstrategic sequential choice model, and a strategic sequential choice model, respectively. Which assumption, and corresponding model, is appropriate depends on the data generating process (DGP) that produced the observations in the sample. Employing a functional form that does not match the DGP is a serious form of model misspecification (Signorino and Yilmaz 2000). In order to be more specific about the alternative statistical models, and to tie the discussion to the study of political phenomena, let us think through a more complicated example. Consider a researcher who has data on actions taken by states in a crisis, where the outcomes are coded as {SQ, Cap2 , W ar}, for the status quo, capitulation, and war, respectively.3 As in our simple example, the appropriate modelling strategy depends the researcher’s assumptions regarding the decision-making process of the states in the crisis. A Word about Notation In the math that follows, we must track the number of alternatives (3), the relevant coefficients (varies from 3 to 6, depending on the model), the number of states (2), and the observation (i). In addition, for each model, we will specify two utility functions, and we must be able to distinguish between the coefficients in each function. In an effort to keep the notation reasonable, we adopt the following conventions: • superscripts always refer to the state, • the observed component of state 1’s utility for the status quo is normalized to 0 (in the independent multinomial probit, the dyad’s utility is normalized to 0), • the observed component of state 2’s utility for capitulation is normalized to 0, 2 The choices in this situation are based on expected utility calculations, but those calculations are made over choices at information sets, not final outcomes. 3 SQ is shorthand for state 1 decides against pressing a claim, and state 2 does nothing. Cap2 is shorthand for state 1 presses a claim, and state 2 capitulates. W ar is shorthand for state 1 presses a claim, and state 2 fights. 3 • the second subscript on β always refers to the utility function the vector is in. The normalizations are substantively innocuous and help ensure an equivalence between the models for the simulation. 2.1 Independent Multinomial Probit or Multinomial Logit The independent multinomial probit or multinomial logit model arises if we assume that two states in a crisis jointly decide on which action to take. The “game tree” for this model is depicted in the first panel of Figure 1. [Figure 1 about here.] Let the three choices the dyad can make, {SQ, Cap2 , W ar}, be denoted 1, 2, and 3, respectively. Let the utility of choice 2 (capitulation) for dyad i be, Yi2∗ = x1i2 β 122 + x1i3 β 132 + x2i3 β 232 + i2 , where the superscript distinguishes between the two states in the dyad, xm ij is a k ≥ 1 row vector of observable variables that explain choice j, for state m, in the ith dyad, and β m j is a k × 1 column vector of unknown and unobservable parameters associated with state m making choice j.4 If we assume utility maximizing behavior, alternative j will be chosen by dyad i if, Yij∗ > Yik∗ , ∀ k = j. Let yij = 1 0 if Yij∗ > Yik∗ , ∀ k = j, otherwise. Assuming that the ij ’s are independent and identically distributed, the independent multinomial probit or multinomial logit is derived by assuming that each has either a normal distribution or an extreme value distribution. Assuming normality and independence across utilities, m βm j has an additional subscript, βj2 , that distinguishes this coefficient vector in the dyad’s utility function for choice 2 from the coefficient vector in the dyad’s utility function for choice 3, βm j3 . 4 Each 4 Pr(yi1 = 1) = = = = Pr(yi2 = 1) Pr(Yi1∗ > Yi2∗ , Yi1∗ > Yi3∗ ) Pr(i1 > x1i2 β 122 + x1i3 β 132 + x2i3 β 232 + i2 , i1 > x1i2 β 123 + x1i3 β 133 + x2i3 β 233 + i3 ) Pr(i1 − i2 > x1i2 β 122 + x1i3 β 132 + x2i3 β 232 , i1 − i3 > x1i2 β 123 + x1i3 β 133 + x2i3 β 233 ) Pr(i2 − i1 < −x1i2 β 122 − x1i3 β 132 − x2i3 β 232 , i3 − i1 < −x1i2 β 123 − x1i3 β 133 − x2i3 β 233 ) = Φ(−x1i2 β 122 − x1i3 β 132 − x2i3 β 232 ) · Φ(−x1i2 β 123 − x1i3 β 133 − x2i3 β 233 ) = Φ(x1i2 β 122 + x1i3 β 132 + x2i3 β 232 ) · Φ(−x1i2 (β 123 − β 122 ) − x1i3 (β 133 − β 132 ) − x2i3 (β 233 − β 232 )) Pr(yi3 = 1) = Φ(x1i2 β 123 + x1i3 β 133 + x2i3 β 233 ) · Φ(−x1i2 (β 122 − β 123 ) − x1i3 (β 132 − β 133 ) − x2i3 (β 232 − β 233 )). Although we have assumed that the errors are normally distributed, we do not get the familiar multinomial probit model as we have also assumed that the errors between the utilities (as opposed to the observations) are uncorrelated. The independent multinomial probit model is therefore equivalent to the multinomial logit model. While the independence assumption is restrictive, we make it because the multinomial logit model has seen significant use in the international relations literature. 2.2 Nonstrategic Sequential Choice Model The nonstrategic sequential choice model arises if we assume that two states in a crisis, a challenger and a defender, make their choices sequentially. The game tree for this model is depicted in the second panel of Figure 1. State 1 chooses either to accept the status quo or to press a claim against state 2 based on its utility for the status quo versus its utility for either capitulation by state 2 or war. If state 1 decides to press a claim, then state 2 must decide whether to capitulate or fight. Let the utility of state 1 not choosing alternative SQ be, 1 1 1 1 2 2 1 Yi1∗ 1̄ = xi2 β 2 + xi3 β 3 + xi3 β 32 + i1̄ , where the superscript refers to the state in question, x1ij is a k ≥ 1 row vector of observable variables that explain state 1’s choice of j, β 1j is a k × 1 column vector of unknown and unobservable parameters associated with a state making choice j, and 1i1̄ is random error associated with state 1’s choice.5 5 β2 has an extra subscript that distinguishes this coefficient vector from the coefficient 32 vector in state 2’s utility function, β233 . 5 If we assume utility maximizing behavior, alternative SQ will be chosen by state 1 if, Yi11∗ > Yi1∗ 1̄ . Let, 1 yij = 1 0 if, for state 1, Yij∗ > Yik∗ , ∀ k = j, otherwise. Similarly, let the utility of alternative W AR for state 2 be, Yi32∗ = x2i3 β 233 + 2i3 , where the observable component of state 2’s utility for CAP2 has been normalized to 0, and the other notation follows the notation used for state 1’s utility. If we assume utility maximizing behavior, alternative W AR will be chosen by state 2 if, Yi32∗ > Yi22∗ . Again, let, 2 yij = 1 0 if, for state 2, Yij∗ > Yik∗ , ∀ k = j, otherwise. Obviously, the functional form of this model is quite different from that of the independent multinomial probit model. Assuming that the errors are drawn from a normal distribution, the relevant probabilities are given by, 1 Pr(yi1 = 1|x1i , x2i ) = 1 − Φ(x1i2 β 12 + x1i3 β 13 + x2i3 β 232 ) 2 1 Pr(yi3 = 0, yi1 = 0|x1i , x2i ) = Φ2 [−x2i3 β 232 , x1i2 β 12 + x1i3 β 13 + x2i3 β 232 , −ρ] 2 1 Pr(yi3 = 1, yi1 = 0|x1i , x2i ) = Φ2 [x2i3 β 232 , x1i2 β 12 + x1i3 β 13 + x2i3 β 232 , ρ], where Φ and Φ2 stand for the cumulative normal and the bivariate cumulative normal, respectively, and ρ is the correlation between the utilities for state 1 and state 2. 2.3 Strategic Sequential Choice Model The strategic sequential choice model arises if we assume that two states in a crisis, a challenger and a defender, make their choices sequentially, and that state 1 conditions its behavior on what it believes state 2 will do. The game tree for this model is depicted in the final panel of Figure 1. State 1 chooses either to accept the status quo or to press a claim against state 2, conditioning on state 1’s belief about whether state 2 will fight or not. If state 1 decides to press a claim, then state 2 decides whether to capitulate or fight. Working backward up the game tree, let the utility of alternative W AR for state 2 be, 6 Yi32∗ = x2i3 β 233 + 2i3 , where the notation is the same as in Section 2.2. If we assume utility maximizing behavior, alternative W AR will be chosen by state 2 if, Yi32∗ > Yi22∗ . Let, 2 yij = 1 0 if, for state 2, Yij∗ > Yik∗ , ∀ k = j, otherwise. Assuming state 2’s decision is independent of the actions of state 16 , the relevant probabilities are, 2 Pr(yi3 = 0|x1i , x2i ) 2 Pr(yi3 = 1|x1i , x2i ) = = 1 − Φ(x2i3 β 233 ) Φ(x2i3 β 233 ). Working up the game tree, the utility of state 1 not choosing alternative SQ is, 2 2 1 1 2 2 1 1 Yi1∗ 1̄ = (1 − Φ(xi3 β 33 ))xi2 β 22 + (Φ(xi3 β 33 ))xi3 β 32 + i1̄ . Letting pi2 = 1 − Φ(x2i3 β 233 ) and pi3 = 1 − pi2 , the relevant probabilities are, 1 Pr(yi1 = 1|x1i , x2i ) 1 Pr(yi1 = 0|x1i , x2i ) = = 1 − Φ(pi2 (x1i2 β 122 ) + pi3 (x1i3 β 132 )) Φ(pi2 (x1i2 β 122 ) + pi3 (x1i3 β 132 )). The strategic choice model differs from the model in Section 2.2 in that state 1’s utility for pressing a claim is conditioned on state 2’s utility for war, and that the errors across the two utilities are uncorrelated. 3 Nonnested Model Testing The three models in Section 2 are nonnested in terms of their functional forms.7 Determining which of these functional forms is closest to the true, but unknown, specification requires, in practice, the use of discrimination tests that are still new to the vast majority of political scientists. Two of the easiest and least controversial of these tests are the Vuong test (Vuong 1989) and a distributionfree test introduced by Clarke (2003b). 6 We make this assumption for mathematical convenience. a definition of nonnested and methods of determining whether rival models are nonnested, see (Clarke 2001). 7 For 7 Both of these tests are based on the Kullback-Leibler information criteria (Kullback and Leibler 1951). Vuong (1989) defines the KLIC as, KLIC ≡ E0 [ln h0 (Yi |Xi )] − E0 [ln f (Yi |Xi ; β∗ )], where h0 (.|.) is the true conditional density of Yi given Xi (that is, the true but unknown model), E0 is the expectation under the true model, and β∗ are the pseudo-true values of β (the estimates of β when f (Yi |Xi ) is not the true model). The best model is the model that minimizes the KLIC, for the best model is the one that is closest to the true specification. We should therefore choose the model that maximizes E0 [ln f (Yi |Xi ; β∗ )]. In other words, one model should be selected over another if the individual log-likelihoods of that model are significantly larger than the individual log-likelihoods of the rival model. 3.1 The Vuong Test The null hypothesis of Vuong’s test is, f (Yt |Xt ; β∗ ) H0 : E0 ln = 0, g(Yt |Zt ; γ∗ ) which states that the two models are equally close to the true specification.8 The expected value in the above hypothesis is unknown. Vuong demonstrates that under fairly general conditions, 1 f (Yt |Xt ; β∗ ) a.s. LRn (β̂n , γ̂n ) −→ E0 ln , n g(Yt |Zt ; γ∗ ) which means that the expected value can be consistently estimated by n1 times the likelihood ratio statistic. The actual test is then, under H0 : LRn (β̂n , γ̂n ) D √ −→ N (0, 1), ( n)ω̂n where LRn (β̂n , γ̂n ) ≡ Lfn (β̂n ) − Lgn (γ̂n ) and ω̂n2 2 n 2 n 1 1 f (Yt |Xt ; β̂n ) f (Yt |Xt ; β̂n ) ≡ − ln . ln n i=1 g(Yt |Zt ; γ̂n ) n i=1 g(Yt |Zt ; γ̂n ) The Vuong test can be described in simple terms. If the null hypothesis is true, the average value of the log-likelihood ratio should be zero. If Hf is true, the average value of the log-likelihood ratio should be significantly greater than zero. If the reverse is true, the average value of the log-likelihood ratio should 8γ ∗ and Zt in model g are analogous to β∗ and Xt in model f . 8 be significantly less than zero. In other words, the Vuong test statistic is simply the average log-likelihood ratio suitably normalized. The log-likelihoods used in the Vuong test are affected if the number of coefficients in the two models being estimated is different, and therefore the test must be corrected for the degrees of freedom. Vuong (1989) suggests using a correction that corresponds to either Akaike’s (1973) information criteria or Schwarz’s (1978) Bayesian information criteria. In the simulations that follow, we use the latter, making the adjusted statistic9 : q p ln n − ln n LR̃n (β̂n , γ̂n ) ≡ LRn (β̂n , γ̂n ) − 2 2 where p and q are the number of estimated coefficients in models f and g, respectively. 3.2 The Distribution-Free Test The Vuong test is not an exact test; it is normally distributed asymptotically. Simulations demonstrate that the relatively small sample sizes used in some international relations research present a problem for the power of the test (Clarke 2003b). A brief explanation will provide some intuition.10 As stated above, under the null hypothesis, the Vuong statistic is distributed as a standard normal, LRn (β̂n , γ̂n ) D √ −→ N (0, 1). ( n)ω̂n An equivalent asymptotic result (Greene 2003) is that the mean log-likelihood ratio converges almost surely to a normal distribution with mean 0 and asymptotic variance ω̂n2 /n, 1 ω̂n2 a.s. LRn (β̂n , γ̂n ) −→ N 0, . n n This convergence, however, is quite slow. For sample sizes under 500, the distribution is highly leptokurtic — very much like a double-exponential distribution. As Lehmann (1986) points out, the sign test is the LMP test for testing θ ≤ 0 against θ > 0 when the sample is drawn from a double-exponential distribution. The sign test therefore seems to be the obvious solution for situations in which only small to modest sample sizes are available. Clarke’s (2003b) distribution-free test applies the paired sign test to the differences in the individual log-likelihoods from two nonnested models.11 Whereas the Vuong test determines whether or not the average log-likelihood ratio is statistically different from zero, the proposed test determines whether or not the 9 Which correction factor is used makes no difference to this analysis. Clarke (2003a) for an in-depth explanation. 11 Recall that the log-likelihood reported by statistical software is the sum of the loglikelihoods for each individual observation. As a log-likelihood ratio is simply the difference between two log-likelihoods, there are n log-likelihood ratios — one for each observation. 10 See 9 median log-likelihood ratio is statistically different from zero. If the models are equally close to the true specification, half the log-likelihood ratios should be greater than zero and half should be less than zero. If model f is “better” than model g, more than half the log-likelihood ratios should be greater than zero. Conversely, if model g is “better” than model f , more than half the log-likelihood ratios should be less than zero. The null hypothesis is therefore: H0 : θ = 0 where θ is the median log-likelihood ratio. One of the great strengths of this procedure is that implementing the test is remarkably simple and can be produced by any mainstream statistical software package using the following algorithm12 : 1. Run model f , saving the individual log-likelihoods. 2. Run model g, saving the individual log-likelihoods. 3. Compute the differences and count the number of positive and negative values. 4. The number of positive differences is distributed binomial(n, p = .5). This test, like the Vuong test, may be affected if the number of coefficients in the two models being estimated is different. Once again, we need a correction for the degrees of freedom. The Schwarz correction is: p q ln n − ln n 2 2 where p and q are the number of estimated coefficients in models f and g, respectively. As we are working with the individual log-likelihood ratios, we cannot apply this correction to the “summed” log-likelihood ratio, as we did for the Vuong test. We can, however, apply the average correction to the individual loglikelihood ratios. That is, we correct the individual log-likelihoods for model f by a factor of: p ln n 2n and the individual log-likelihoods for model g by a factor of: q ln n. 2n 12 In what follows, steps 1 and 2 are in the process of being implemented by STATA. Step 3 already exists in STATA because we are making use of the paired sign test. The command is simply “signtest ll1 = ll2 ” where lli are the individual log-likelihoods from one model. 10 4 Monte Carlo Simulation We performed a suite of simulations to determine whether or not the tests from the last section can distinguish between alternative choice models that are nonnested in terms of their functional forms. The simulations also allow us to gauge under what conditions we can expect either the Vuong test or the distribution-free test to have greater relative power. The experiments are based on the discrimination of the three competing choice models discussed in Section 2. In the simulations, we controlled for sample size, data generating process, and signal-to-noise ratio (the variance of the systematic portion of the model versus the variance of the error term). 90 simulations were run, each with 2000 replications, with the following parameters: 1. Data generating process (all coefficients set to 1) (a) strategic (b) selection (c) independent MNP 2. Sample size: 50, 100, 200, 500, 1000, 2000 3. Signal-to-noise ratio: 0.333, 0.5, 1.0, 2.0, 3.0. For each experiment, we calculated both the Vuong and distribution-free test statistics for the following comparisons: 1. strategic v. selection, 2. strategic v. independent MNP, 3. selection v. independent MNP. The observed utilities in the simulation are specified as in Section 2 and Figure 2. In each model, as previously noted, state 1’s observed utility for the status quo is normalized to 0, and state 2’s observed utility for capitulation is normalized to 0.13 [Figure 2 about here.] 13 In the case of the multinomial model, the dyad’s observed utility for the status quo is normalized to 0. 11 4.1 Results Discussion of the simulation results raises two interesting issues. First, given the design of the simulations, we cannot discuss the size of the tests because the null hypothesis is false in every experiment (the models are never equally close to the true DGP). Rejecting the null hypothesis when it is true is therefore not possible. We can, however, discuss the power of the tests in both the correct direction (toward the DGP) and the wrong direction (away from the DGP). Second, we are comparing a continuous test statistic, the Vuong, with a discrete test statistic, the number of positive differences. The problem with this comparison is that for any finite number of observations the exact significance level of the discrete test statistic is unlikely to match the nominal significance level selected for the simulation.14 Absent identical exact significance levels, power comparisons may be quite misleading (Gibbons and Chakraborti 1992). We employ a randomized decision rule to solve this problem (Lehmann 1986). Let a test statistic τ be in the rejection region with probability 1 if τ ≥ c2 and with probability ρ if c1 ≤ τ ≤ c2 . The nominal significance level can therefore be achieved, even with a small-n discrete test statistic, using the following rule: Pr(τ ≥ c1 |H0 ) + ρ · Pr(c1 ≤ τ ≤ c2 |H0 ) = α where c1 < c2 . The power levels we report, therefore, are for equivalent nominal significance levels. The results of the simulation for the discrimination of the strategic model against the selection model are in Table 1. The table shows the number of correct decisions out of 2000 for each simulation. As expected, both sample size and signalto-noise ratio affect the probability of the tests choosing the correct model. The power of both tests increases with sample size and as the signal-to-noise ratio moves from low to high. [Table 1 about here.] The main conclusion we can draw from Table 1 is that discrimination between strategic choice models and nonstrategic selection models is quite feasible. Under the very worst conditions, a small sample size and a very low signal-tonoise ratio, both tests chose the correct model in more than 50% of the repetitions. Both tests are, of course, consistent, and both tests consequently chose the correct model with probability 1 as the sample size approached 2000 and the signal-to-noise ratio approached 3. Table 1 also shows that the distribution-free test outperformed the Vuong test in 23 out of 30 simulations. The greater relative power of the distributionfree test does not, however, come without a price. While the Vuong test never chose the incorrect model, it either chose the correct model or chose neither, the distribution-free test did occasionally choose the incorrect model. Table 2 14 A discrete test statistic has a limited number of probabilities (the number of “jump points” in the CDF) that can serve as α. These probabilities are exact significance levels. 12 shows the percentage of replications in which the distribution-free test chose the incorrect model. Even under the worst conditions, the test chose the incorrect model in only 0.8% of the replications. The benefits gained from the greater power of the distribution-free test clearly outweighs the minuscule probability of rejecting the null in favor of the incorrect model. [Table 2 about here.] The strategic model was tested, not only against the selection model, also against the independent MNP.15 Similarly, the independent MNP was tested against the selection model. In the first of these simulations, the strategic against the independent MNP, both tests chose the correct model nearly 100% of the time. In the second, the independent MNP against the selection, both tests chose the selection model over the independent MNP nearly 100% of the time. The power of the tests in these simulations is due to the “distance” of the selection model and the independent MNP from the strategic model. Simulations of the Kullback-Leibler distance show that the selection model is roughly half the distance of the independent MNP from the strategic model.16 The tests therefore have no trouble choosing the strategic model over the independent MNP and the selection model over the independent MNP. 4.2 Discussion The results indicate the conditions under which we can expect either the Vuong test or the distribution-free test to perform well. The distribution-free test outperforms the Vuong test under the most difficult conditions: small sample sizes and low signal-to-noise ratio. That is, the distribution-free test outperforms the Vuong test in situations in which correct discrimination is least likely to occur. This result makes sense as the sign test has greatest relative power under conditions where power tends to be smallest (Bradley 1968). The simulation results should be of great interest to substantive scholars. The results are important in that small sample studies, though not the majority, are common in international relations research. For example, seven recent smalln studies in conflict studies are Huth (1988), which has an n of 58; Huth, Gelpi, and Bennett (1993), which has an n of 97; Reiter and Stam (1998), which has an n of 197; Signorino and Tarar (2002), which has an n of 58; Bennett and Stam (1996), which has an n of 169; Benoit (1996), which has an n of 97; and Pollins (1996), which has an n of 161. In addition, the current state of international relations theory suggests that signal-to-noise ratio in our models is probably low. As the simulations demonstrate, a low signal-to-noise ratio makes discrimination more difficult. A test that works under these conditions is surely welcome. 15 Full 16 See results are available from the authors upon request. Clarke 2001 for the procedure for estimating the KLIC by simulation. 13 5 Conclusion The purpose of this paper is to combine and extend the research programmes of its authors. We have provided a framework in which it is possible to compare strategic models against one another, or against nonstrategic alternatives. At the same time, we have extended nonnested model testing in political science to situations where the rival models are nonnested in terms of their functional forms. We have demonstrated that discriminating between strategic choice models and various alternative nonstrategic choice models is feasible even under adverse conditions. While the distribution-free test has greater relative power in situations where power is needed, both tests perform well and are extremely easy to implement. There is no reason why a substantively-oriented scholar should need to simply assume a functional form whether or not that functional form is strategic. We hope that future scholars will use these results to move the scientific study of international relations beyond its current state. 14 References Achen, Christopher. 1982. Interpreting and Using Regression. Beverly Hills: Sage. Akaike, H. 1973. “Information Theory and an Extension of the Likelihood Ratio Principle.” In Second International Symposium of Information Theory, eds. B.N. Petrov and F. Csaki. Minnesota Studies in the Philosophy of Science, Budapest: Akademinai Kiado. Bennett, D. Scott, and Allan C. Stam. 1996. “The Duration of Interstate Wars, 1816-1985.” American Political Science Review 90:239–257. Benoit, Kenneth. 1996. “Democracies Really Are More Pacific (in General): Reexamining Regime Type and War Involvement.” Journal of Conflict Resolution 40:636–657. Bradley, James V. 1968. Distribution-Free Statistical Tests. New Jersey: Prentice-Hall. Clarke, Kevin A. 2001. “Testing Nonnested Models of International Relations: Reevaluating Realism.” American Journal of Political Science 45:724–744. Clarke, Kevin A. 2003a. “A Distribution-Free Discrimination Test.” Clarke, Kevin A. 2003b. “Nonparametric Model Discrimination in International Relations.” Journal of Conflict Resolution 47:72–93. Gibbons, Jean Dickinson, and Subhabrata Chakraborti. 1992. Nonparametric Statistical Inference. 3 ed. New York: Marcel Dekker, Inc. Greene, William H. 2003. Econometric Analysis. 5 ed. New Jersey: Prentice Hall. Huth, Paul, Christopher Gelpi, and D. Scott Bennett. 1993. “The Escalation of Great Power Militarized Disputes: Testing Rational Deterrence Theory and Structural Realism.” American Political Science Review 87:609–623. Huth, Paul K. 1988. Extended Deterrence and the Prevention of War . New Haven: Yale University Press. Kullback, Solomon, and R.A. Leibler. 1951. “On Information and Sufficiency.” Annals of Mathematical Statistics 22:79–86. Lehmann, E. L. 1986. Testing Statistical Hypotheses. 2 ed. New York: John Wiley. Pollins, Brian M. 1996. “Global Political Order, Economic Change, and Armed Conflict: Coevolving Systems and the Use of Force.” American Political Science Review 90:103–117. Reiter, Dan, and Allan Stam. 1998. “Democracy, War Initiation, and Victory.” American Political Science Review 92:377–389. 15 Schwarz, G. 1978. “Estimating the Dimension of a Model.” Annals of Statistics 6:461–464. Signorino, Curtis S. 1999. “Strategic Interaction and the Statistical Analysis of International Conflict.” American Political Science Review 93:411–433. Signorino, Curtis S., and Ahmer Tarar. 2002. “A Unified Theory and Test of Extended Immediate Deterrence.” Signorino, Curtis S., and Kuzey Yilmaz. 2000. “Strategic Misspecification in Discrete Choice Models.” Unpublished manuscript. University of Rochester. Vuong, Quang. 1989. “Likelihood ratio tests for model selection and nonnested hypotheses.” Econometrica 57:307–333. 16 List of Figures 1 2 Alternative Nonnested Choice Models . . . . . . . . . . . . . . . Observed Utilities for the Simulation . . . . . . . . . . . . . . . . 17 18 19 Figure 1: Alternative Nonnested Choice Models 1,2 1 SQ Cap2 W ar SQ (2.1) Multinomial Model 2 SQ Cap2 W ar 2 Cap2 W ar (2.2) Nonstrategic Sequential Selection 1 (2.3) Strategic Model 18 Figure 2: Observed Utilities for the Simulation 1,2 1 1 1 1 1 2 2 1 yi1∗ 1̄ = xi2 β2 + xi3 β3 + xi3 β32 + i1̄ SQ Cap2 1 1 xi2 β22 + 1 + x1i3 β32 2 2 xi3 β32 W ar 1 x1i2 β23 + SQ 0 2 x2i3 β33 1 1 1 1 1 yi1∗ 1̄ = ρ2 · xi2 β22 + ρ3 · xi3 β32 + i1̄ 2 SQ 2∗ 2 yi3 = x2i3 β33 + 2i3 0 Cap2 W ar 1 x1i3 β32 1 x1i2 β22 0 Cap2 W ar (2.2) Nonstrategic Sequential Selection 1 2 2∗ 2 yi3 = x2i3 β33 + 2i3 1 + x1i3 β33 (2.1) Multinomial Model 2 x2i3 β33 (2.3) Strategic Model 19 List of Tables 1 2 Monte Carlo Comparison: Strategic Choice v. Selection . . . . . Monte Carlo Comparison: Strategic Choice v. Selection . . . . . 20 21 22 Table 1: Monte Carlo Comparison: Strategic Choice v. Selection Signal to Noise Ratio Size Test 0.333 0.5 1.0 2.0 3.0 50 Clarke Vuong 1287 1125 1424 1155 1670 1309 1918 1624 1988 1935 100 Clarke Vuong 1467 1304 1592 1322 1826 1550 1977 1875 1995 1970 200 Clarke Vuong 1586 1468 1733 1536 1921 1790 1993 1979 2000 2000 500 Clarke Vuong 1733 1681 1863 1819 1988 1990 1999 2000 2000 2000 1000 Clarke Vuong 1846 1857 1916 1955 1998 2000 2000 2000 2000 2000 2000 Clarke Vuong 1925 1967 1985 1999 2000 2000 2000 2000 2000 2000 21 Table 2: Monte Carlo Comparison: Strategic Choice v. Selection Signal to Noise Ratio Size 0.333 0.5 1.0 2.0 3.0 50 0.008 0.0065 0.0025 0.0005 0.0 100 0.006 0.007 0.0005 0.0 0.0 200 0.0065 0.0035 0.0005 0.0 0.0 500 0.0055 0.0015 0.0 0.0 0.0 1000 0.0015 0.001 0.0 0.0 0.0 2000 0.0 0.0 0.0 0.0 0.0 22
© Copyright 2026 Paperzz