Nonnested Tests for Strategic Choice Models

Discriminating Methods:
Nonnested Tests for Strategic Choice Models
Kevin A. Clarke
(585) 275-5217
[email protected]
†
Curt S. Signorino
(585) 273-4760
[email protected]
Department of Political Science
University of Rochester
Harkness Hall
Rochester, NY 14627-0146
(585) 271-1616 (F)
Abstract
We consider a researcher attempting to choose, on the available data, between a strategic choice model and various nonstrategic choice models.
Such a researcher is faced with choosing between rival models that are
nonnested in terms of their functional forms. We discuss both a parametric and distribution-free procedure for making this choice, and demonstrate through a monte carlo simulation that discrimination is possible.
The results of the simulation also allow us to compare the relative power
of the two tests.
May 4, 2003
† An earlier version of this paper was presented at the annual meeting of the International
Studies Association in Portland, OR (2003); we thank the participants for their comments.
Support from the National Science Foundation (Grant SES-0213771) is gratefully acknowledged.
1
Introduction
The empirical study of political science has, in recent years, undergone something akin to a sea-change. Where it was once common for political scientists to
employ a linear functional form regardless of the theory being tested, we now see
a new attention being paid to the connection between theory and model. One
result of this attention has been the realization that a disconnect between theory and model results in serious misspecification and parameter bias (Signorino
and Yilmaz 2000).
Now that researchers are encouraged to derive a functional form directly
from the theoretical model under consideration, new challenges have arisen.
Although Signorino (1999) has demonstrated that traditional specifications of
statistical models are generally inconsistent with strategic theories of political
science, no rigorous framework has emerged for comparing strategic models
against one another, or against nonstrategic models. While it is clear that
strategic specifications provide different answers from traditional specifications,
it not yet clear that these strategic specifications are, in fact, superior. The
question at hand is whether or not strategic specifications are “closer” than
nonstrategic models to the data generating processes at work in the field of
international relations.
Compounding the difficulty of the issue is the fact that not all theory is
detailed enough to allow the derivation of a functional form suitable for testing.
Achen (1982), for example, argues that good social science theory is functionally non-specific. An empirical researcher faced with choosing a specification
under these conditions needs guidance in choosing between the many functional
forms, some strategic and some not, that may be used to model most political
phenomena.
In either case, empirical researchers need tools that allow models with different functional forms to be compared. Models with different functional forms,
however, are generally nonnested (neither model is a special case of the other
model). Discriminating between nonnested models requires specialized tests
that are rarely used in political science research. Clarke (2001) introduced the
issue of nonnested testing to political science, and Clarke (2003b) introduced a
simple distribution-free test for nonnested model discrimination. These articles,
however, consider only models that are nonnested in terms of their covariates.
Testing models that are nonnested in terms of their functional forms is a natural
extension of this line of research.
This paper, in essence, represents a first step in combining the research
programmes of its authors. Our hope is that future scholars will be able to
develop statistical models that are truly consistent with their theories and be
able to test those theories against one another. The attainment of these goals
will allow the scientific study of international relations to move significantly
beyond its current state.
The paper proceeds as follows. In Section 2, we detail a discrete choice problem that may be modelled in a number of different ways. We demonstrate that
the assumptions made by the researcher regarding the decision-making process
determine the correct functional form. In Section 3, we review both a parametric and distribution-free discrimination test for nonnested models. Finally, in
Section 4, we perform a suite of monte carlo simulations to demonstrate that
discriminating between alternative choice models characterized by nonnested
functional forms is feasible. The simulations also allow us to compare the relative power of the two tests.
2
Alternative Choice Models
Consider a researcher who wishes to model the decision-making process of two
individuals choosing between three unordered alternatives. A common approach
to this problem is to associate choice with utility and assume that individual i
receives utility Uij when alternative j is chosen. There exist, however, at least
three common “random utility” specifications the researcher could use to model
this particular choice problem.
As a way of fixing ideas, let us think through a simple example in which a
married couple is attempting to choose between three different modes of transportation for getting to work.1 The options include riding their bikes, taking
a bus, or taking a taxi. A simple, but often overlooked, point is that how this
example should be modelled depends on how the couple in question makes their
decision.
For instance, the researcher might assume that the couple makes their choice
jointly. That is, the couple decides, as a unit, which transportation option
provides them the greatest utility. The researcher should therefore model the
situation as two individuals maximizing their joint utility and choosing option
∗
∗
> Uik
, ∀ k = j.
j if Uij
Choosing a mode of transportation as a unit is not, however, the only method
two individuals could employ to make this decision. Another possibility is that
the researcher could assume that the choice occurs sequentially. Individual
1 chooses between option j and not option j, leaving individual 2 to choose
between options k and l. Returning to our example, the wife may choose between
riding bikes or not based on her expected utility for these options, but not take
into account her husband’s preferences regarding the bus versus a taxi. If the
wife decides that the couple will not ride their bikes to work, her husband then
may choose between taking a bus or taking a taxi based on his expected utility.
Note the censoring we may observe given this arrangement. If the wife chooses
bike riding, we do not learn anything about the choice the husband might have
made.
A third possibility is that the researcher might assume that the choice occurs
not only sequentially, but strategically. That is, individual 1 chooses between
option j and not option j, based not on an expected utility calculation as in
the example above, but conditioning on his or her belief about individual 2’s
1 We
will consider a more complex political example in the next section.
2
subsequent choice between k and l.2 Continuing our running example, the wife
decides whether or not the couple should ride their bikes, but makes this choice
taking into consideration whether she believes her husband will choose a taxi
over a bus or vice versa. The husband then makes his choice between the bus
or a taxi based on a straightforward expected utility calculation. This decisionmaking process is quite different from the last example where the wife does not
take into account what she thinks her husband will do. As in the last example,
if the wife chooses bike riding, we do not observe the husband’s choice.
Were we to write down a statistical model for each of these situations,
the three different assumptions made by the researcher regarding the couple’s
decision-making process would lead to three different functional forms: the independent multinomial probit or logit, a nonstrategic sequential choice model,
and a strategic sequential choice model, respectively. Which assumption, and
corresponding model, is appropriate depends on the data generating process
(DGP) that produced the observations in the sample. Employing a functional
form that does not match the DGP is a serious form of model misspecification
(Signorino and Yilmaz 2000).
In order to be more specific about the alternative statistical models, and to tie
the discussion to the study of political phenomena, let us think through a more
complicated example. Consider a researcher who has data on actions taken by
states in a crisis, where the outcomes are coded as {SQ, Cap2 , W ar}, for the
status quo, capitulation, and war, respectively.3 As in our simple example, the
appropriate modelling strategy depends the researcher’s assumptions regarding
the decision-making process of the states in the crisis.
A Word about Notation
In the math that follows, we must track the number of alternatives (3), the
relevant coefficients (varies from 3 to 6, depending on the model), the number
of states (2), and the observation (i). In addition, for each model, we will
specify two utility functions, and we must be able to distinguish between the
coefficients in each function. In an effort to keep the notation reasonable, we
adopt the following conventions:
• superscripts always refer to the state,
• the observed component of state 1’s utility for the status quo is normalized to 0 (in the independent multinomial probit, the dyad’s utility is
normalized to 0),
• the observed component of state 2’s utility for capitulation is normalized
to 0,
2 The choices in this situation are based on expected utility calculations, but those calculations are made over choices at information sets, not final outcomes.
3 SQ is shorthand for state 1 decides against pressing a claim, and state 2 does nothing.
Cap2 is shorthand for state 1 presses a claim, and state 2 capitulates. W ar is shorthand for
state 1 presses a claim, and state 2 fights.
3
• the second subscript on β always refers to the utility function the vector
is in.
The normalizations are substantively innocuous and help ensure an equivalence between the models for the simulation.
2.1
Independent Multinomial Probit or Multinomial Logit
The independent multinomial probit or multinomial logit model arises if we
assume that two states in a crisis jointly decide on which action to take. The
“game tree” for this model is depicted in the first panel of Figure 1.
[Figure 1 about here.]
Let the three choices the dyad can make, {SQ, Cap2 , W ar}, be denoted 1, 2,
and 3, respectively. Let the utility of choice 2 (capitulation) for dyad i be,
Yi2∗ = x1i2 β 122 + x1i3 β 132 + x2i3 β 232 + i2 ,
where the superscript distinguishes between the two states in the dyad, xm
ij is
a k ≥ 1 row vector of observable variables that explain choice j, for state m, in
the ith dyad, and β m
j is a k × 1 column vector of unknown and unobservable
parameters associated with state m making choice j.4
If we assume utility maximizing behavior, alternative j will be chosen by
dyad i if,
Yij∗ > Yik∗ , ∀ k = j.
Let
yij =
1
0
if Yij∗ > Yik∗ , ∀ k = j,
otherwise.
Assuming that the ij ’s are independent and identically distributed, the
independent multinomial probit or multinomial logit is derived by assuming
that each has either a normal distribution or an extreme value distribution.
Assuming normality and independence across utilities,
m
βm
j has an additional subscript, βj2 , that distinguishes this coefficient vector in the
dyad’s utility function for choice 2 from the coefficient vector in the dyad’s utility function
for choice 3, βm
j3 .
4 Each
4
Pr(yi1 = 1)
=
=
=
=
Pr(yi2 = 1)
Pr(Yi1∗ > Yi2∗ , Yi1∗ > Yi3∗ )
Pr(i1 > x1i2 β 122 + x1i3 β 132 + x2i3 β 232 + i2 ,
i1 > x1i2 β 123 + x1i3 β 133 + x2i3 β 233 + i3 )
Pr(i1 − i2 > x1i2 β 122 + x1i3 β 132 + x2i3 β 232 ,
i1 − i3 > x1i2 β 123 + x1i3 β 133 + x2i3 β 233 )
Pr(i2 − i1 < −x1i2 β 122 − x1i3 β 132 − x2i3 β 232 ,
i3 − i1 < −x1i2 β 123 − x1i3 β 133 − x2i3 β 233 )
=
Φ(−x1i2 β 122 − x1i3 β 132 − x2i3 β 232 ) · Φ(−x1i2 β 123 − x1i3 β 133 − x2i3 β 233 )
=
Φ(x1i2 β 122 + x1i3 β 132 + x2i3 β 232 ) ·
Φ(−x1i2 (β 123 − β 122 ) − x1i3 (β 133 − β 132 ) − x2i3 (β 233 − β 232 ))
Pr(yi3 = 1)
=
Φ(x1i2 β 123 + x1i3 β 133 + x2i3 β 233 ) ·
Φ(−x1i2 (β 122 − β 123 ) − x1i3 (β 132 − β 133 ) − x2i3 (β 232 − β 233 )).
Although we have assumed that the errors are normally distributed, we do
not get the familiar multinomial probit model as we have also assumed that
the errors between the utilities (as opposed to the observations) are uncorrelated. The independent multinomial probit model is therefore equivalent to
the multinomial logit model. While the independence assumption is restrictive,
we make it because the multinomial logit model has seen significant use in the
international relations literature.
2.2
Nonstrategic Sequential Choice Model
The nonstrategic sequential choice model arises if we assume that two states in
a crisis, a challenger and a defender, make their choices sequentially. The game
tree for this model is depicted in the second panel of Figure 1. State 1 chooses
either to accept the status quo or to press a claim against state 2 based on its
utility for the status quo versus its utility for either capitulation by state 2 or
war. If state 1 decides to press a claim, then state 2 must decide whether to
capitulate or fight.
Let the utility of state 1 not choosing alternative SQ be,
1 1
1 1
2 2
1
Yi1∗
1̄ = xi2 β 2 + xi3 β 3 + xi3 β 32 + i1̄ ,
where the superscript refers to the state in question, x1ij is a k ≥ 1 row vector
of observable variables that explain state 1’s choice of j, β 1j is a k × 1 column
vector of unknown and unobservable parameters associated with a state making
choice j, and 1i1̄ is random error associated with state 1’s choice.5
5 β2 has an extra subscript that distinguishes this coefficient vector from the coefficient
32
vector in state 2’s utility function, β233 .
5
If we assume utility maximizing behavior, alternative SQ will be chosen by
state 1 if,
Yi11∗ > Yi1∗
1̄ .
Let,
1
yij
=
1
0
if, for state 1, Yij∗ > Yik∗ , ∀ k = j,
otherwise.
Similarly, let the utility of alternative W AR for state 2 be,
Yi32∗ = x2i3 β 233 + 2i3 ,
where the observable component of state 2’s utility for CAP2 has been normalized to 0, and the other notation follows the notation used for state 1’s utility.
If we assume utility maximizing behavior, alternative W AR will be chosen by
state 2 if,
Yi32∗ > Yi22∗ .
Again, let,
2
yij
=
1
0
if, for state 2, Yij∗ > Yik∗ , ∀ k = j,
otherwise.
Obviously, the functional form of this model is quite different from that of
the independent multinomial probit model. Assuming that the errors are drawn
from a normal distribution, the relevant probabilities are given by,
1
Pr(yi1
= 1|x1i , x2i ) = 1 − Φ(x1i2 β 12 + x1i3 β 13 + x2i3 β 232 )
2
1
Pr(yi3
= 0, yi1
= 0|x1i , x2i ) = Φ2 [−x2i3 β 232 , x1i2 β 12 + x1i3 β 13 + x2i3 β 232 , −ρ]
2
1
Pr(yi3
= 1, yi1
= 0|x1i , x2i ) = Φ2 [x2i3 β 232 , x1i2 β 12 + x1i3 β 13 + x2i3 β 232 , ρ],
where Φ and Φ2 stand for the cumulative normal and the bivariate cumulative
normal, respectively, and ρ is the correlation between the utilities for state 1
and state 2.
2.3
Strategic Sequential Choice Model
The strategic sequential choice model arises if we assume that two states in a
crisis, a challenger and a defender, make their choices sequentially, and that
state 1 conditions its behavior on what it believes state 2 will do. The game
tree for this model is depicted in the final panel of Figure 1. State 1 chooses
either to accept the status quo or to press a claim against state 2, conditioning
on state 1’s belief about whether state 2 will fight or not. If state 1 decides
to press a claim, then state 2 decides whether to capitulate or fight. Working
backward up the game tree, let the utility of alternative W AR for state 2 be,
6
Yi32∗ = x2i3 β 233 + 2i3 ,
where the notation is the same as in Section 2.2.
If we assume utility maximizing behavior, alternative W AR will be chosen
by state 2 if,
Yi32∗ > Yi22∗ .
Let,
2
yij
=
1
0
if, for state 2, Yij∗ > Yik∗ , ∀ k = j,
otherwise.
Assuming state 2’s decision is independent of the actions of state 16 , the
relevant probabilities are,
2
Pr(yi3
= 0|x1i , x2i )
2
Pr(yi3 = 1|x1i , x2i )
=
=
1 − Φ(x2i3 β 233 )
Φ(x2i3 β 233 ).
Working up the game tree, the utility of state 1 not choosing alternative SQ
is,
2 2
1 1
2 2
1 1
Yi1∗
1̄ = (1 − Φ(xi3 β 33 ))xi2 β 22 + (Φ(xi3 β 33 ))xi3 β 32 + i1̄ .
Letting pi2 = 1 − Φ(x2i3 β 233 ) and pi3 = 1 − pi2 , the relevant probabilities are,
1
Pr(yi1
= 1|x1i , x2i )
1
Pr(yi1 = 0|x1i , x2i )
=
=
1 − Φ(pi2 (x1i2 β 122 ) + pi3 (x1i3 β 132 ))
Φ(pi2 (x1i2 β 122 ) + pi3 (x1i3 β 132 )).
The strategic choice model differs from the model in Section 2.2 in that state
1’s utility for pressing a claim is conditioned on state 2’s utility for war, and
that the errors across the two utilities are uncorrelated.
3
Nonnested Model Testing
The three models in Section 2 are nonnested in terms of their functional forms.7
Determining which of these functional forms is closest to the true, but unknown,
specification requires, in practice, the use of discrimination tests that are still
new to the vast majority of political scientists. Two of the easiest and least
controversial of these tests are the Vuong test (Vuong 1989) and a distributionfree test introduced by Clarke (2003b).
6 We
make this assumption for mathematical convenience.
a definition of nonnested and methods of determining whether rival models are
nonnested, see (Clarke 2001).
7 For
7
Both of these tests are based on the Kullback-Leibler information criteria
(Kullback and Leibler 1951). Vuong (1989) defines the KLIC as,
KLIC ≡ E0 [ln h0 (Yi |Xi )] − E0 [ln f (Yi |Xi ; β∗ )],
where h0 (.|.) is the true conditional density of Yi given Xi (that is, the true
but unknown model), E0 is the expectation under the true model, and β∗ are
the pseudo-true values of β (the estimates of β when f (Yi |Xi ) is not the true
model). The best model is the model that minimizes the KLIC, for the best
model is the one that is closest to the true specification. We should therefore
choose the model that maximizes E0 [ln f (Yi |Xi ; β∗ )]. In other words, one model
should be selected over another if the individual log-likelihoods of that model
are significantly larger than the individual log-likelihoods of the rival model.
3.1
The Vuong Test
The null hypothesis of Vuong’s test is,
f (Yt |Xt ; β∗ )
H0 : E0 ln
= 0,
g(Yt |Zt ; γ∗ )
which states that the two models are equally close to the true specification.8
The expected value in the above hypothesis is unknown. Vuong demonstrates
that under fairly general conditions,
1
f (Yt |Xt ; β∗ )
a.s.
LRn (β̂n , γ̂n ) −→ E0 ln
,
n
g(Yt |Zt ; γ∗ )
which means that the expected value can be consistently estimated by n1 times
the likelihood ratio statistic. The actual test is then,
under H0 :
LRn (β̂n , γ̂n ) D
√
−→ N (0, 1),
( n)ω̂n
where
LRn (β̂n , γ̂n ) ≡ Lfn (β̂n ) − Lgn (γ̂n )
and
ω̂n2
2 n
2
n
1
1 f (Yt |Xt ; β̂n )
f (Yt |Xt ; β̂n )
≡
−
ln
.
ln
n i=1
g(Yt |Zt ; γ̂n )
n i=1
g(Yt |Zt ; γ̂n )
The Vuong test can be described in simple terms. If the null hypothesis is
true, the average value of the log-likelihood ratio should be zero. If Hf is true,
the average value of the log-likelihood ratio should be significantly greater than
zero. If the reverse is true, the average value of the log-likelihood ratio should
8γ
∗
and Zt in model g are analogous to β∗ and Xt in model f .
8
be significantly less than zero. In other words, the Vuong test statistic is simply
the average log-likelihood ratio suitably normalized.
The log-likelihoods used in the Vuong test are affected if the number of
coefficients in the two models being estimated is different, and therefore the
test must be corrected for the degrees of freedom. Vuong (1989) suggests using
a correction that corresponds to either Akaike’s (1973) information criteria or
Schwarz’s (1978) Bayesian information criteria. In the simulations that follow,
we use the latter, making the adjusted statistic9 :
q
p ln n −
ln n
LR̃n (β̂n , γ̂n ) ≡ LRn (β̂n , γ̂n ) −
2
2
where p and q are the number of estimated coefficients in models f and g,
respectively.
3.2
The Distribution-Free Test
The Vuong test is not an exact test; it is normally distributed asymptotically.
Simulations demonstrate that the relatively small sample sizes used in some
international relations research present a problem for the power of the test
(Clarke 2003b). A brief explanation will provide some intuition.10 As stated
above, under the null hypothesis, the Vuong statistic is distributed as a standard
normal,
LRn (β̂n , γ̂n ) D
√
−→ N (0, 1).
( n)ω̂n
An equivalent asymptotic result (Greene 2003) is that the mean log-likelihood
ratio converges almost surely to a normal distribution with mean 0 and asymptotic variance ω̂n2 /n,
1
ω̂n2
a.s.
LRn (β̂n , γ̂n ) −→ N 0,
.
n
n
This convergence, however, is quite slow. For sample sizes under 500, the
distribution is highly leptokurtic — very much like a double-exponential distribution. As Lehmann (1986) points out, the sign test is the LMP test for testing
θ ≤ 0 against θ > 0 when the sample is drawn from a double-exponential distribution. The sign test therefore seems to be the obvious solution for situations
in which only small to modest sample sizes are available.
Clarke’s (2003b) distribution-free test applies the paired sign test to the differences in the individual log-likelihoods from two nonnested models.11 Whereas
the Vuong test determines whether or not the average log-likelihood ratio is statistically different from zero, the proposed test determines whether or not the
9 Which
correction factor is used makes no difference to this analysis.
Clarke (2003a) for an in-depth explanation.
11 Recall that the log-likelihood reported by statistical software is the sum of the loglikelihoods for each individual observation. As a log-likelihood ratio is simply the difference
between two log-likelihoods, there are n log-likelihood ratios — one for each observation.
10 See
9
median log-likelihood ratio is statistically different from zero. If the models
are equally close to the true specification, half the log-likelihood ratios should
be greater than zero and half should be less than zero. If model f is “better” than model g, more than half the log-likelihood ratios should be greater
than zero. Conversely, if model g is “better” than model f , more than half the
log-likelihood ratios should be less than zero. The null hypothesis is therefore:
H0 : θ = 0
where θ is the median log-likelihood ratio.
One of the great strengths of this procedure is that implementing the test is
remarkably simple and can be produced by any mainstream statistical software
package using the following algorithm12 :
1. Run model f , saving the individual log-likelihoods.
2. Run model g, saving the individual log-likelihoods.
3. Compute the differences and count the number of positive and negative
values.
4. The number of positive differences is distributed binomial(n, p = .5).
This test, like the Vuong test, may be affected if the number of coefficients
in the two models being estimated is different. Once again, we need a correction
for the degrees of freedom. The Schwarz correction is:
p q
ln n −
ln n
2
2
where p and q are the number of estimated coefficients in models f and g,
respectively.
As we are working with the individual log-likelihood ratios, we cannot apply
this correction to the “summed” log-likelihood ratio, as we did for the Vuong
test. We can, however, apply the average correction to the individual loglikelihood ratios. That is, we correct the individual log-likelihoods for model f
by a factor of:
p ln n
2n
and the individual log-likelihoods for model g by a factor of:
q ln n.
2n
12 In what follows, steps 1 and 2 are in the process of being implemented by STATA. Step
3 already exists in STATA because we are making use of the paired sign test. The command
is simply “signtest ll1 = ll2 ” where lli are the individual log-likelihoods from one model.
10
4
Monte Carlo Simulation
We performed a suite of simulations to determine whether or not the tests
from the last section can distinguish between alternative choice models that
are nonnested in terms of their functional forms. The simulations also allow
us to gauge under what conditions we can expect either the Vuong test or the
distribution-free test to have greater relative power.
The experiments are based on the discrimination of the three competing
choice models discussed in Section 2. In the simulations, we controlled for
sample size, data generating process, and signal-to-noise ratio (the variance of
the systematic portion of the model versus the variance of the error term).
90 simulations were run, each with 2000 replications, with the following
parameters:
1. Data generating process (all coefficients set to 1)
(a) strategic
(b) selection
(c) independent MNP
2. Sample size: 50, 100, 200, 500, 1000, 2000
3. Signal-to-noise ratio: 0.333, 0.5, 1.0, 2.0, 3.0.
For each experiment, we calculated both the Vuong and distribution-free
test statistics for the following comparisons:
1. strategic v. selection,
2. strategic v. independent MNP,
3. selection v. independent MNP.
The observed utilities in the simulation are specified as in Section 2 and
Figure 2. In each model, as previously noted, state 1’s observed utility for the
status quo is normalized to 0, and state 2’s observed utility for capitulation is
normalized to 0.13
[Figure 2 about here.]
13 In the case of the multinomial model, the dyad’s observed utility for the status quo is
normalized to 0.
11
4.1
Results
Discussion of the simulation results raises two interesting issues. First, given
the design of the simulations, we cannot discuss the size of the tests because the
null hypothesis is false in every experiment (the models are never equally close
to the true DGP). Rejecting the null hypothesis when it is true is therefore not
possible. We can, however, discuss the power of the tests in both the correct
direction (toward the DGP) and the wrong direction (away from the DGP).
Second, we are comparing a continuous test statistic, the Vuong, with a
discrete test statistic, the number of positive differences. The problem with this
comparison is that for any finite number of observations the exact significance
level of the discrete test statistic is unlikely to match the nominal significance
level selected for the simulation.14 Absent identical exact significance levels,
power comparisons may be quite misleading (Gibbons and Chakraborti 1992).
We employ a randomized decision rule to solve this problem (Lehmann 1986).
Let a test statistic τ be in the rejection region with probability 1 if τ ≥ c2 and
with probability ρ if c1 ≤ τ ≤ c2 . The nominal significance level can therefore
be achieved, even with a small-n discrete test statistic, using the following rule:
Pr(τ ≥ c1 |H0 ) + ρ · Pr(c1 ≤ τ ≤ c2 |H0 ) = α
where c1 < c2 . The power levels we report, therefore, are for equivalent nominal
significance levels.
The results of the simulation for the discrimination of the strategic model against
the selection model are in Table 1. The table shows the number of correct decisions out of 2000 for each simulation. As expected, both sample size and signalto-noise ratio affect the probability of the tests choosing the correct model. The
power of both tests increases with sample size and as the signal-to-noise ratio
moves from low to high.
[Table 1 about here.]
The main conclusion we can draw from Table 1 is that discrimination between strategic choice models and nonstrategic selection models is quite feasible.
Under the very worst conditions, a small sample size and a very low signal-tonoise ratio, both tests chose the correct model in more than 50% of the repetitions. Both tests are, of course, consistent, and both tests consequently chose
the correct model with probability 1 as the sample size approached 2000 and
the signal-to-noise ratio approached 3.
Table 1 also shows that the distribution-free test outperformed the Vuong
test in 23 out of 30 simulations. The greater relative power of the distributionfree test does not, however, come without a price. While the Vuong test never
chose the incorrect model, it either chose the correct model or chose neither,
the distribution-free test did occasionally choose the incorrect model. Table 2
14 A discrete test statistic has a limited number of probabilities (the number of “jump points”
in the CDF) that can serve as α. These probabilities are exact significance levels.
12
shows the percentage of replications in which the distribution-free test chose the
incorrect model. Even under the worst conditions, the test chose the incorrect
model in only 0.8% of the replications. The benefits gained from the greater
power of the distribution-free test clearly outweighs the minuscule probability
of rejecting the null in favor of the incorrect model.
[Table 2 about here.]
The strategic model was tested, not only against the selection model, also
against the independent MNP.15 Similarly, the independent MNP was tested
against the selection model. In the first of these simulations, the strategic
against the independent MNP, both tests chose the correct model nearly 100%
of the time. In the second, the independent MNP against the selection, both
tests chose the selection model over the independent MNP nearly 100% of the
time. The power of the tests in these simulations is due to the “distance” of the
selection model and the independent MNP from the strategic model. Simulations of the Kullback-Leibler distance show that the selection model is roughly
half the distance of the independent MNP from the strategic model.16 The
tests therefore have no trouble choosing the strategic model over the independent MNP and the selection model over the independent MNP.
4.2
Discussion
The results indicate the conditions under which we can expect either the Vuong
test or the distribution-free test to perform well. The distribution-free test
outperforms the Vuong test under the most difficult conditions: small sample
sizes and low signal-to-noise ratio. That is, the distribution-free test outperforms
the Vuong test in situations in which correct discrimination is least likely to
occur. This result makes sense as the sign test has greatest relative power
under conditions where power tends to be smallest (Bradley 1968).
The simulation results should be of great interest to substantive scholars.
The results are important in that small sample studies, though not the majority,
are common in international relations research. For example, seven recent smalln studies in conflict studies are Huth (1988), which has an n of 58; Huth, Gelpi,
and Bennett (1993), which has an n of 97; Reiter and Stam (1998), which has
an n of 197; Signorino and Tarar (2002), which has an n of 58; Bennett and
Stam (1996), which has an n of 169; Benoit (1996), which has an n of 97; and
Pollins (1996), which has an n of 161.
In addition, the current state of international relations theory suggests that
signal-to-noise ratio in our models is probably low. As the simulations demonstrate, a low signal-to-noise ratio makes discrimination more difficult. A test
that works under these conditions is surely welcome.
15 Full
16 See
results are available from the authors upon request.
Clarke 2001 for the procedure for estimating the KLIC by simulation.
13
5
Conclusion
The purpose of this paper is to combine and extend the research programmes
of its authors. We have provided a framework in which it is possible to compare
strategic models against one another, or against nonstrategic alternatives. At
the same time, we have extended nonnested model testing in political science
to situations where the rival models are nonnested in terms of their functional
forms.
We have demonstrated that discriminating between strategic choice models
and various alternative nonstrategic choice models is feasible even under adverse
conditions. While the distribution-free test has greater relative power in situations where power is needed, both tests perform well and are extremely easy
to implement. There is no reason why a substantively-oriented scholar should
need to simply assume a functional form whether or not that functional form
is strategic. We hope that future scholars will use these results to move the
scientific study of international relations beyond its current state.
14
References
Achen, Christopher. 1982. Interpreting and Using Regression. Beverly Hills:
Sage.
Akaike, H. 1973. “Information Theory and an Extension of the Likelihood
Ratio Principle.” In Second International Symposium of Information
Theory, eds. B.N. Petrov and F. Csaki. Minnesota Studies in the Philosophy of Science, Budapest: Akademinai Kiado.
Bennett, D. Scott, and Allan C. Stam. 1996. “The Duration of Interstate
Wars, 1816-1985.” American Political Science Review 90:239–257.
Benoit, Kenneth. 1996. “Democracies Really Are More Pacific (in General):
Reexamining Regime Type and War Involvement.” Journal of Conflict
Resolution 40:636–657.
Bradley, James V. 1968. Distribution-Free Statistical Tests. New Jersey:
Prentice-Hall.
Clarke, Kevin A. 2001. “Testing Nonnested Models of International Relations: Reevaluating Realism.” American Journal of Political Science
45:724–744.
Clarke, Kevin A. 2003a. “A Distribution-Free Discrimination Test.”
Clarke, Kevin A. 2003b. “Nonparametric Model Discrimination in International Relations.” Journal of Conflict Resolution 47:72–93.
Gibbons, Jean Dickinson, and Subhabrata Chakraborti. 1992. Nonparametric Statistical Inference. 3 ed. New York: Marcel Dekker, Inc.
Greene, William H. 2003. Econometric Analysis. 5 ed. New Jersey: Prentice
Hall.
Huth, Paul, Christopher Gelpi, and D. Scott Bennett. 1993. “The Escalation of Great Power Militarized Disputes: Testing Rational Deterrence
Theory and Structural Realism.” American Political Science Review
87:609–623.
Huth, Paul K. 1988. Extended Deterrence and the Prevention of War . New
Haven: Yale University Press.
Kullback, Solomon, and R.A. Leibler. 1951. “On Information and Sufficiency.” Annals of Mathematical Statistics 22:79–86.
Lehmann, E. L. 1986. Testing Statistical Hypotheses. 2 ed. New York: John
Wiley.
Pollins, Brian M. 1996. “Global Political Order, Economic Change, and
Armed Conflict: Coevolving Systems and the Use of Force.” American
Political Science Review 90:103–117.
Reiter, Dan, and Allan Stam. 1998. “Democracy, War Initiation, and Victory.” American Political Science Review 92:377–389.
15
Schwarz, G. 1978. “Estimating the Dimension of a Model.” Annals of Statistics 6:461–464.
Signorino, Curtis S. 1999. “Strategic Interaction and the Statistical Analysis of International Conflict.” American Political Science Review
93:411–433.
Signorino, Curtis S., and Ahmer Tarar. 2002. “A Unified Theory and Test
of Extended Immediate Deterrence.”
Signorino, Curtis S., and Kuzey Yilmaz. 2000. “Strategic Misspecification in Discrete Choice Models.” Unpublished manuscript. University
of Rochester.
Vuong, Quang. 1989. “Likelihood ratio tests for model selection and nonnested hypotheses.” Econometrica 57:307–333.
16
List of Figures
1
2
Alternative Nonnested Choice Models . . . . . . . . . . . . . . .
Observed Utilities for the Simulation . . . . . . . . . . . . . . . .
17
18
19
Figure 1: Alternative Nonnested Choice Models
1,2
1
SQ
Cap2
W ar
SQ
(2.1) Multinomial Model
2
SQ
Cap2
W ar
2
Cap2
W ar
(2.2) Nonstrategic Sequential Selection
1
(2.3) Strategic Model
18
Figure 2: Observed Utilities for the Simulation
1,2
1
1 1
1 1
2 2
1
yi1∗
1̄ = xi2 β2 + xi3 β3 + xi3 β32 + i1̄
SQ
Cap2
1 1
xi2 β22 +
1
+
x1i3 β32
2 2
xi3 β32
W ar
1
x1i2 β23
+
SQ
0
2
x2i3 β33
1 1
1 1
1
yi1∗
1̄ = ρ2 · xi2 β22 + ρ3 · xi3 β32 + i1̄
2
SQ
2∗
2
yi3
= x2i3 β33
+ 2i3
0
Cap2
W ar
1
x1i3 β32
1
x1i2 β22
0
Cap2
W ar
(2.2) Nonstrategic Sequential Selection
1
2
2∗
2
yi3
= x2i3 β33
+ 2i3
1
+
x1i3 β33
(2.1) Multinomial Model
2
x2i3 β33
(2.3) Strategic Model
19
List of Tables
1
2
Monte Carlo Comparison: Strategic Choice v. Selection . . . . .
Monte Carlo Comparison: Strategic Choice v. Selection . . . . .
20
21
22
Table 1: Monte Carlo Comparison: Strategic Choice v. Selection
Signal to Noise Ratio
Size
Test
0.333
0.5
1.0
2.0
3.0
50
Clarke
Vuong
1287
1125
1424
1155
1670
1309
1918
1624
1988
1935
100
Clarke
Vuong
1467
1304
1592
1322
1826
1550
1977
1875
1995
1970
200
Clarke
Vuong
1586
1468
1733
1536
1921
1790
1993
1979
2000
2000
500
Clarke
Vuong
1733
1681
1863
1819
1988
1990
1999
2000
2000
2000
1000
Clarke
Vuong
1846
1857
1916
1955
1998
2000
2000
2000
2000
2000
2000
Clarke
Vuong
1925
1967
1985
1999
2000
2000
2000
2000
2000
2000
21
Table 2: Monte Carlo Comparison: Strategic Choice v. Selection
Signal to Noise Ratio
Size
0.333
0.5
1.0
2.0
3.0
50
0.008
0.0065
0.0025
0.0005
0.0
100
0.006
0.007
0.0005
0.0
0.0
200
0.0065
0.0035
0.0005
0.0
0.0
500
0.0055
0.0015
0.0
0.0
0.0
1000
0.0015
0.001
0.0
0.0
0.0
2000
0.0
0.0
0.0
0.0
0.0
22