A Weighting Approach to Causal Effects and Additive Interaction in

American Journal of Epidemiology
ª The Author 2011. Published by Oxford University Press on behalf of the Johns Hopkins Bloomberg School of
Public Health. All rights reserved. For permissions, please e-mail: [email protected].
Vol. 174, No. 10
DOI: 10.1093/aje/kwr334
Advance Access publication:
October 19, 2011
Practice of Epidemiology
A Weighting Approach to Causal Effects and Additive Interaction in Case-Control
Studies: Marginal Structural Linear Odds Models
Tyler J. VanderWeele* and Stijn Vansteelandt
* Correspondence to Dr. Tyler J. VanderWeele, Departments of Epidemiology and Biostatistics, Harvard School of Public Health,
677 Huntington Avenue, Boston, MA 02115 (e-mail: [email protected]).
Initially submitted October 28, 2010; accepted for publication June 13, 2011.
Estimates of additive interaction from case-control data are often obtained by logistic regression; such models can
also be used to adjust for covariates. This approach to estimating additive interaction has come under some criticism
because of possible misspecification of the logistic model: If the underlying model is linear, the logistic model will be
misspecified. The authors propose an inverse probability of treatment weighting approach to causal effects and
additive interaction in case-control studies. Under the assumption of no unmeasured confounding, the approach
amounts to fitting a marginal structural linear odds model. The approach allows for the estimation of measures
of additive interaction between dichotomous exposures, such as the relative excess risk due to interaction, using
case-control data without having to rely on modeling assumptions for the outcome conditional on the exposures
and covariates. Rather than using conditional models for the outcome, models are instead specified for the exposures
conditional on the covariates. The approach is illustrated by assessing additive interaction between genetic and
environmental factors using data from a case-control study.
case-control studies; interaction; linear model; structural model; synergism; weighting
Abbreviations: OR, odds ratio; RERI, relative excessive risk due to interaction; RR, risk ratio.
In this paper, we consider the use of inverse probability
weighting to estimate causal effects in unmatched case-control
studies. The approach we take effectively amounts to fitting
what may be defined as a marginal structural linear odds model
to case-control data. The approach is quite general. However,
the methodological development here was motivated by the
problem of attempting to assess interaction on the additive
scale by using data from a case-control study.
Additive interaction is often assessed by estimating a quantity sometimes referred to as the ‘‘relative excess risk due to
interaction’’ (RERI) (1). If there are 2 dichotomous factors
(call them A and B) and we let RRij denote the risk ratio
(RR) comparing A ¼ i and B ¼ j with A ¼ B ¼ 0, then the
relative excess risk due to interaction for the risk ratio is
defined by RR11 – RR10 – RR01 þ 1. In a case-control study
with ‘‘cumulative design’’ where controls are sampled from
among those disease free at the end of follow-up, the odds
ratio will generally be used to estimate the effect of the
factors A and B. If the outcome is rare, then the risk ratios
can be approximated by odds ratios (ORs). If we let ORij denote the odds ratio comparing A ¼ i and B ¼ j with A ¼ B ¼ 0,
then the relative excess risk due to interaction for the odds
ratio is defined by
RERI ¼ OR11 OR10 OR01 þ 1:
Throughout this paper, we will be using RERI on the odds
ratio scale. The RERI can be used to give a measure of interaction on the additive scale for case-control data when the
outcome is rare.
Measures of interaction on the additive scale, such as RERI,
are generally what is thought to be of most importance in
considering public health implications (2–5). This is because
additive interaction allows one to assess whether an intervention would have a larger absolute effect in one subpopulation
versus another. Estimates of RERI are also useful in detecting synergism between 2 factors within the sufficient cause
framework (5, 6). Under the assumption that both exposures
1197
Am J Epidemiol. 2011;174(10):1197–1203
1198 VanderWeele and Vansteelandt
have neutral or causative effects for all individuals (i.e., the
effects are positive monotonic), RERI > 0 implies such
synergism (5–7); without such monotonicity assumptions,
one can still test for synergism by testing RERI > 1 (6–8).
Several articles (9–13) have considered the problem of
estimating and providing confidence intervals for RERI.
Earlier work by Hosmer and Lemeshow (9) and Assmann
et al. (10) considered logistic regression models and the
delta method or bootstrapping for confidence intervals for
RERI. More recent work, such as that by Richardson and
Kaufman (12), has considered using a linear odds model
(5, 12–14), namely, odds ¼ exp(b0)(1 þ b1A þ b2B þ b3AB),
to obtain estimates and confidence intervals for RERI. Under
the linear odds model, the coefficient and confidence interval for b3 can be interpreted as RERI and its confidence
interval (12). Both the approach using logistic regression and
the approach using the linear odds model can be used to give
confidence intervals even when controlling for other covariates, provided that the regression models are correctly
specified. When there are covariates for which adjustments
are to be made, these can likewise be included in either a logistic regression or linear odds model. In these cases, however,
the relation between the covariates and the outcome needs
to be correctly specified; failure to do so can lead to invalid
inferences (14, 15). Here, we provide an alternative approach
to estimating the RERI using weighting that will be applicable
irrespective of whether the true underlying model relating the
outcome and the covariates is a logistic model, a linear odds
model, or some other model. This weighting approach will
require that models for the exposures are correctly specified
but will not require correctly specifying a model for the
outcome.
A WEIGHTING APPROACH TO RERI
Suppose that data come from an unmatched case-control
study (we make a few remarks about how the approach might
be adapted to matched case-control designs in the Web
Appendix, which is posted on the Journal’s Web site (http://
aje.oupjournals.org/)), where controls are sampled from those
who are disease free at the end of follow-up. Suppose also that
there are covariates for which control is to be made. Finally,
we suppose that the outcome is rare; this assumption will be
needed to estimate the weights and so that the relative excess
risk due to interaction estimate for the odds ratio can genuinely
be interpreted as a measure of additive interaction.
We show in the Web Appendix that, instead of including
covariates in a linear odds model or a logistic model, one can
use a weighting approach for covariate adjustment as follows.
One first estimates inverse probability of treatment weights
using data just on the controls. This could be done by using
2 logistic regressions between the controls: 1) a logistic regression of the first exposure on the covariates and 2) a logistic
regression of the second exposure on the covariates and the
first exposure. Again, both regressions use data only on the
controls.
For each individual, a weight is obtained for the first
exposure (call it wA) by taking the inverse of the predicted
probability from the first logistic regression that the individual
had the exposure level of A that was in fact present. Likewise,
a weight is obtained for the second exposure (call it wB) by
taking the inverse of the predicted probability from the second logistic regression that the individual had the exposure
level of B that was in fact present. These are referred to as
inverse-probability-of-treatment weights (16). Multiplying
these 2 weights together (wA 3 wB) gives the overall weight
for the individual. Note that, although the logistic regression
models are fit for the controls only, the predicted probabilities
and weights are calculated for each individual in the sample
(both cases and controls).
If a linear odds model conditional on the 2 exposures,
Odds ¼ expðb0 Þð1 þ b1 A þ b2 B þ b3 ABÞ;
is fit to case-control data with these weights, then, under a rare
outcome assumption, the coefficient and confidence interval
for b3 in this weighted linear odds regression will give an
estimate of relative excess risk due to interaction for the
standardized odds ratios adjusted for the covariates. Further
detail is given in Appendix 1 and justification in the Web
Appendix. Even though the procedure above requires estimation of the weights, if robust standard errors are used,
it will still yield confidence intervals and estimates of the
standard error that are conservative, as in other weighting
approaches (16).
As discussed in Appendix 1, the procedure of using the
controls to calculate the weights and then fitting weighted
linear odds models is applicable to the estimation of causal
effects more generally and not simply to the RERI. Under
the rare outcome assumption and provided that the set of
covariates for which adjustment is made suffice to control
for confounding, the procedure essentially amounts to fitting
what can be defined as a ‘‘marginal structural linear odds
model.’’ The approach also extends discussion of marginal
structural models for interaction in cohort studies (17) to
case-control studies. SAS implementation is also given in
Appendix 2.
ILLUSTRATION
The approach is illustrated with data from a case-control
study of lung cancer at Massachusetts General Hospital (18)
of 1,836 cases and 1,452 controls. Eligible cases included any
person over the age of 18 years, with a diagnosis of primary
lung cancer that was further confirmed by a lung pathologist. The controls were recruited from among the friends or
spouses of cancer patients or the friends or spouses of other
surgery patients in the same hospital. Potential controls that
carried a previous diagnosis of any cancer (other than nonmelanoma skin cancer) were excluded from participation.
The study included information on smoking and genotype
information on locus 15q25.1.
Genetic variants on 15q25.1 have been found to be associated with both smoking and lung cancer (19–21). We examine
whether there is interaction on the additive scale between
the effects of smoking (ever vs. never) and the genetic variant
(0 vs. 1/2 T alleles at rs8034191). Covariate data include age
(continuous), gender, and educational history (college degree
Am J Epidemiol. 2011;174(10):1197–1203
Weighting, Causal Effects, and Additive Interaction
or more, yes/no). Analyses were limited to Caucasians. Using
the procedure above gives an estimate of RERI ¼ 2.71
(percentile bootstrap confidence interval: 1.68, 4.04) with
2,000 bootstrapped samples. The estimate suggests positive
interaction on the additive scale between smoking and the
genetic variant. The estimate and confidence interval suggest
synergism in the sufficient cause sense (5–8) even without
assumptions about monotonicity.
SIMULATION
Data were simulated by using the sample sizes and datagenerating mechanism observed in the aforementioned casecontrol study of lung cancer. Specifically, in each of 1,000
simulation experiments, a large number of binary measurements for gender and educational history were drawn from
their empirical multinomial distribution, and normally distributed age measurements were drawn conditionally on those.
Smoking status and genotype were generated under logistic
regression models, allowing for a dependence between both
exposures, conditional on covariates. Lung cancer status was
also generated from a logistic model, allowing for main effects
and an interaction between both exposures, as well as the main
effects of all 3 covariates. Subsequently, data were retained
for 1,836 cases and 1,452 controls. The code generating the
simulated data is given in the Web Appendix.
Five simulation experiments were constructed, corresponding to 5 different population outcome means. Data were
analyzed by 1) the proposed approach using logistic regression
models for both exposures and also by 2) a conditional linear
odds regression:
OddsðYÞ ¼ expðc0 þ c4 CÞð1 þ c1 a þ c2 b þ c3 abÞ:
Note that this model is deliberately misspecified under our
data-generating model, in view of the aforementioned concerns
about correctly specifying conditional linear odds regression
models. Note further also that the weighting approach proposed in this paper is subject to model misspecification bias
because it involves equating a model for both exposures
correctly specified for the population with the corresponding
model in controls. Reported coverage is for 95% Wald confidence intervals based on conservative standard errors for
1199
the proposed procedure and 95% likelihood intervals (12) for
the conditional linear odds approach.
The results are summarized in Table 1. They confirm the
adequate performance of the proposed approach at disease
prevalences below 10%. As predicted by the theory, the confidence intervals based on the proposed approach are conservative (this could be addressed through the use of bootstrap
confidence intervals). The performance of both approaches
deteriorates with larger disease prevalences, because the
aforementioned degree of model misspecification grows with
increasing prevalence.
We note that Månsson et al. (22) also evaluated an inverseprobability weighting approach to total causal effects in the
context of propensity score analysis using simulations and
found reasonably good coverage probabilities for the approach
under some, but not all, scenarios. Unfortunately, they did not
directly report the outcome prevalence for their simulations,
so it is difficult to compare their simulation scenarios with
those reported here.
DISCUSSION AND IMPLICATIONS FOR MODEL
SPECIFICATION
Motivation for this weighting approach to testing and
estimation for additive interaction comes in part from a paper by Skrondal (14). Skrondal cautioned against the use of
RERI and logistic regression to examine interaction on the
additive scale in case-control studies. He noted that, if the
underlying linear risk model with covariates was correctly
specified, the additive interaction then took on a single value
in all strata of the covariates but RERI would vary across
strata (what he called the ‘‘uniqueness problem’’). Although
this is true, it is nevertheless still the case that, if the linear
risk model with covariates is correctly specified, RERI will
be of the same sign in all strata of the covariates, and thus
RERI > 0 (the condition for synergism under monotonicity)
will either hold in all strata of covariates or in no stratum.
The condition RERI > 1 (i.e., the condition for synergism
without monotonicity) could vary across strata, but it is
also the case that without monotonicity, synergism may be
present in some strata of the covariates but not in others,
even if the underlying linear risk model with covariates is
correct (6, 7).
Table 1. Results for Simulation Comparing Marginal Structural Linear Odds Models and Conditional Linear Odds Models at Different Outcome
Prevalences When Both Are Potentially Subject to Misspecification
Marginal Structural Linear Odds Model
E (Y ),%
Empirical SE
Estimated SE
Conditional Linear Odds Model
Bias
Empirical SE
Coverage
Probability, %
2.55
0.069
0.54
94.1
2.50
0.12
0.53
94.8
2.38
0.17
0.51
94.5
98.1
2.09
0.49
0.50
85.0
31.7
1.54
1.05
0.52
41.3
Coverage
Probability, %
RERI
Bias
0.5
2.53
0.061
0.56
0.71
96.1
1.3
2.46
0.14
0.55
0.71
95.7
3.3
2.29
0.23
0.52
0.65
98.4
8.3
2.29
0.25
0.52
0.62
18.3
1.35
1.20
0.55
0.55
RERI
Abbreviations: E(Y ), outcome prevalence; RERI, relative excess risk due to interaction; SE, standard error.
Am J Epidemiol. 2011;174(10):1197–1203
1200 VanderWeele and Vansteelandt
Perhaps more importantly, Skrondal noted that, if the linear
risk model with covariates was correctly specified, the logistic
model with covariates would not be correctly specified (what
he called the ‘‘misspecification problem’’). Greenland (23)
further remarked that parsimonious logistic regression models
impose nonadditivity. By estimating RERI with such models,
one thus risks introducing a bias toward nonadditivity on the
additive scale. Skrondal noted that the linear odds model with
covariates could be used to identify the parameters of the
linear risk model up to a constant of proportionality (although
this, in fact, only holds under an assumption that the outcome
is rare). It is, however, conversely also the case that, if in fact
the logistic regression model with covariates is correctly
specified, then the linear risk model and the linear odds
model will not be. The advantage of using the weighting approach we described above is that it will give valid estimates
of the relative excess risk due to interaction for the standardized odds ratios irrespective of whether the true underlying
model is a linear risk model with covariates, a logistic model
with covariates, or some other model, provided that the models
used for the weights are correctly specified. The weighting
approach described above does not require a correctly specified
conditional model for the outcome given the exposures and
covariates, but it does require a correctly specified model for
the weights; it thus effectively transfers the problem of correctly specifying a model from the outcome to the exposures,
as is also the case with marginal structural models in cohort
studies (16, 17). This weighting approach that avoids having
to correctly specify an outcome model conditional on the
covariates is furthermore desirable within the context of testing
for synergism because an outcome model imposes certain
assumptions within the sufficient cause framework, whereas
a model for the exposures does not (17). In related work, we
are developing a doubly robust approach that will yield valid
inferences for unmatched case-control data if either a conditional model for the exposures or a conditional model
for the outcome is correctly specified. The use of Bayesian
approaches may also be of interest (24).
ACKNOWLEDGMENTS
Author affiliations: Department of Epidemiology, Harvard
School of Public Health, Boston, Massachusetts (Tyler J.
VanderWeele); Department of Biostatistics, Harvard
School of Public Health, Boston, Massachusetts (Tyler J.
VanderWeele); and Department of Applied Mathematics and
Computer Science, Ghent University, Ghent, Belgium (Stijn
Vansteelandt).
The research was supported by grants ES017876 and
HD060696 from the National Institutes of Health.
Conflict of interest: none declared.
REFERENCES
1. Rothman KJ. Modern Epidemiology. 1st ed. Boston, MA:
Little, Brown and Company; 1986.
2. Blot WJ, Day NE. Synergism and interaction: are they
equivalent? Am J Epidemiol. 1979;100(1):99–100.
3. Rothman KJ, Greenland S, Walker AM. Concepts of interaction.
Am J Epidemiol. 1980;112(4):467–470.
4. Saracci R. Interaction and synergism. Am J Epidemiol. 1980;
112(4):465–466.
5. Rothman KJ, Greenland S, Lash TL. Modern Epidemiology.
3rd ed. Philadelphia, PA: Lippincott Williams & Wilkins;
2008.
6. VanderWeele TJ, Robins JM. The identification of synergism
in the sufficient-component-cause framework. Epidemiology.
2007;18(3):329–339.
7. VanderWeele TJ. Sufficient cause interactions and statistical
interactions. Epidemiology. 2009;20(1):6–13.
8. VanderWeele TJ. Empirical tests for compositional epistasis
[letter]. Nat Rev Genet. 2010;11(2):166.
9. Hosmer DW, Lemeshow S. Confidence interval estimation of
interaction. Epidemiology. 1992;3(5):452–456.
10. Assmann SF, Hosmer DW, Lemeshow S, et al. Confidence
intervals for measures of interaction. Epidemiology. 1996;
7(3):286–290.
11. Nie L, Chu H, Li F, et al. Relative excess risk due to interaction: resampling-based confidence intervals. Epidemiology.
2010;21(4):552–556.
12. Richardson DB, Kaufman JS. Estimation of the relative excess
risk due to interaction and associated confidence bounds. Am
J Epidemiol. 2009;169(6):756–760.
13. Kuss O, Schmidt-Pokrzywniak A, Stang A. Confidence intervals for the interaction contrast ratio. Epidemiology. 2010;
21(2):273–274.
14. Skrondal A. Interaction as departure from additivity in casecontrol studies: a cautionary note. Am J Epidemiol. 2003;
158(3):251–258.
15. Greenland S. Additive risk versus additive relative risk
models. Epidemiology. 1993;4(1):32–36.
16. Robins JM, Hernán MA, Brumback B. Marginal structural
models and causal inference in epidemiology. Epidemiology.
2000;11(5):550–560.
17. VanderWeele TJ, Vansteelandt S, Robins JM. Marginal structural models for sufficient cause interactions. Am J Epidemiol.
2010;171(4):506–514.
18. Miller DP, Liu G, De Vivo I, et al. Combinations of the variant
genotypes of GSTP1, GSTM1, and p53 are associated with
an increased lung cancer risk. Cancer Res. 2002;62(10):
2819–2823.
19. Hung RJ, McKay JD, Gaborieau V, et al. A susceptibility
locus for lung cancer maps to nicotinic acetylcholine receptor subunit genes on 15q25. Nature. 2008;452(1787):
633–637.
20. Amos CI, Wu X, Broderick P, et al. Genome-wide association
scan of tag SNPs identifies a susceptibility locus for lung
cancer at 15q25.1. Nat Genet. 2008;40(5):616–622.
21. Thorgeirsson TE, Geller F, Sulem P, et al. A variant associated with nicotine dependence, lung cancer and
peripheral arterial disease. Nature. 2008;452(7187):
638–642.
22. Månsson R, Joffe MM, Sun W, et al. On the estimation and use
of propensity scores in case-control and case-cohort studies.
Am J Epidemiol. 2007;166(3):332–339.
23. Greenland S. Interactions in epidemiology: relevance,
identification, and estimation. Epidemiology. 2009;20(1):
14–17.
24. Chu H, Nie L, Cole SR. Estimating the relative excess risk
due to interaction: a Bayesian approach. Epidemiology. 2011;
22(2):242–248.
Am J Epidemiol. 2011;174(10):1197–1203
Weighting, Causal Effects, and Additive Interaction
APPENDIX 1
Marginal Structural Linear Odds Models
Let Y denote the outcome of interest, and let T denote the
exposure(s) of interest. The variable T may include more
than one exposure. In the application to RERI, T consists of
2 exposures (A, B). We let Yt denote the counterfactual value
of Y if T had been set to level t. The odds for Yt can then be
defined as
OddsðYt Þ ¼ PðYt ¼ 1Þ=f1 PðYt ¼ 1Þg:
1201
OddsðYða;bÞ oddsðYð0;0Þ ¼ ðh0 þ h1 a þ h2 b þ h3 ab h0
¼ ðkc0 þ kc1 a
þ kc2 b þ kc3 ab kc0
¼ 1 þ ðc1 c0 Þa
þ ðc2 c0 Þb þ ðc3 c0 Þab:
If, instead, we parameterized the marginal structural linear
odds model as
OddsðYða;bÞ Þ ¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ
A marginal structural linear odds model takes the form:
and fit a weighted conditional linear odds regression,
OddsðYt Þ ¼ gðtÞ#h;
OddsðYÞ ¼ expðc0 Þð1 þ c1 a þ c2 b þ c3 abÞ;
where g(t) is a vector-valued function of t, and h is a vector
of parameters. For example, for 2 exposures A and B, we
would have that T ¼ (A, B), g(t) ¼ g(a, b) ¼ (1, a, b, ab)#,
and h ¼ (h0, h1, h2, h3)# so that the marginal structural linear
odds model is
then the parameter estimates, (c1, c2, c3), from the reparameterized weighted linear odds regression will directly
give consistent estimates of the causal odds ratio parameters
(b1, b2, b3). Note that, to obtain consistent estimates of the
causal odds ratios, both assumption 1 of rare outcome and
assumption 2 of no confounding conditional on C were required. If only assumption 1 holds (i.e., if assumption 2 does
not hold so there is still confounding), the weighting procedure
above will still give consistent estimates for the standardized
odds ratios adjusted for C, namely:
P
PðY ¼ 1j A ¼ a; B ¼ b; C ¼ cÞPðC ¼ cÞ
c
P
1 PðY ¼ 1j A ¼ a; B ¼ b; C ¼ cÞPðC ¼ cÞ
OddsðYða;bÞ Þ ¼ h0 þ h1 a þ h2 b þ h3 ab:
The parameters of the marginal structural linear odds model
can be fit with case-control data by using inverse probability
of treatment weighting if 1) a rare outcome assumption
holds and 2) the set of covariates C suffices to control for
confounding for the effects of T on Y (in counterfactual
notation, this is Yt is independent of T conditional on C).
If so, inverse probability of treatment weights for individual
i is given by wi ¼ 1/P(T ¼ tijC ¼ ci, Y ¼ 0), where the
denominator probability can be estimated by regression
models (multiple regression models if T is multivariate).
For example, for 2 exposures, the weights would be as
follows:
wi ¼
1
1
;
3
PðA ¼ ai jC ¼ ci Þ
PðB ¼ bi jA ¼ ai ; C ¼ ci Þ
where each of the 2 denominator probabilities could be estimated by logistic regression. Under assumptions 1) and
2) fitting a conditional linear odds regression, odds(Y) ¼ g(t)# c
with the observed case-control data, where each subject is
weighted by wi, will give consistent estimates of the corresponding parameters, h, of the marginal structural model up
to a constant of proportionality, k, so that kc ¼ h. For
example, with 2 exposures, A and B, we have that fitting the
conditional linear odds regression:
OddsðYÞ ¼ c0 þ c1 a þ c2 b þ c3 ab;
where each subject weighted by wi will give consistent
estimates of the corresponding parameters of the marginal
structural model (h0, h1, h2, h3) up to a constant of proportionality, that is (kc0, kc1, kc2, kc3) ¼ (h0, h1, h2, h3). Note
that, from the estimates of (c0, c1, c2, c3), we can estimate
the causal odds ratios:
Am J Epidemiol. 2011;174(10):1197–1203
c
¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ:
In the case of estimating RERI with 2 dichotomous factors,
the model specifications
OddsðYða;bÞ Þ ¼ h0 þ h1 a þ h2 b þ h3 ab
and
OddsðYða;bÞ Þ ¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ
are both saturated. Model misspecification is not an issue.
In the more general case, the marginal structural model,
odds(Yt) ¼ g(t)#h, may not be saturated. In this case, the
marginal structural model must be correctly specified to get
valid estimates of causal effects. Moreover, when the marginal
structural model is not saturated, more efficient estimates
of the coefficients can sometimes be obtained by using what
are sometimes called ‘‘stabilized weights’’ (16) rather than
the weights wi given above. Using stabilized weights, one
would weight the conditional linear odds regression by
swi ¼ PðT ¼ ti jY ¼ 0Þ PðT ¼ ti jC ¼ ci ; Y ¼ 0Þ
rather than by
wi ¼ 1 PðT ¼ ti jC ¼ ci ; Y ¼ 0Þ:
1202 VanderWeele and Vansteelandt
APPENDIX 2
SAS Implementation for a Weighting Approach to RERI
We describe how the weighting approach to RERI given above for an unmatched case-control design can be implemented
in SAS statistical software (SAS Institute, Inc., Cary, North Carolina). Suppose the data are stored in a data set defined as
‘‘mydata’’ and the outcome variable is ‘‘y,’’ the 2 binary exposures are ‘‘A’’ and ‘‘B’’ and the covariates are c1, c2, c3. The
following SAS code will create the predicted probabilities above (pa and pb), the 2 weights (wa and wb), the overall inverse
probability weight (ipw), and the estimate of the RERI (b3).
data controldata;
set mydata;
if y¼0;
run;
proc logistic data¼controldata descending outest¼params1;
model A¼ c1 c2 c3;
run;
proc logistic data¼mydata descending inest¼params1;
model A¼ c1 c2 c3/ MAXITER¼0;
output out¼mydata predicted¼pa;
run;
proc logistic data¼controldata descending outest¼params2;
model B ¼ A c1 c2 c3;
run;
proc logistic data¼mydata descending inest¼params2;
model B ¼ A c1 c2 c3/ MAXITER¼0;
output out¼mydata predicted¼pb;
run;
data mydata;
set mydata;
if A¼1 then wa¼1/pa;
if A¼0 then wa¼1/(1-pa);
if B¼1 then wb¼1/pb;
if B¼0 then wb¼1/(1-pb);
ipw¼wa*wb;
run;
proc nlmixed data¼mydata;
odds¼exp(b0)*(1þb1*Aþb2*Bþb3*A*B);
model y ~ binary(odds/(1þodds));
replicate ipw;
run;
The final procedure of this code (proc nlmixed) is essentially that given by Richardson and Kaufman (12) plus an additional
replicate statement to weight the observations. The coefficient b3 produced by SAS will be an estimate of the RERI.
Unfortunately, SAS proc nlmixed does not have an option to produce robust standard errors that are needed for a valid
confidence interval in this case. Here, we provide additional SAS code to give bootstrapped standard errors for the estimate
obtained from the SAS code above. In the Web Appendix, we also give alternative SAS code using the delta method to give
robust standard errors.
The following SAS code can be used to obtain percentile bootstrap confidence intervals for this estimate. The code uses
2,000 samples, but users could increase this to a larger number as well by changing the second line of the code.
data boot;
do sample¼1 to 2000;
do i¼1 to nobs;
pt¼round(ranuni(12)*nobs);
set mydata nobs¼nobs point¼pt ;
output;
end;
end;
stop;
run;
Am J Epidemiol. 2011;174(10):1197–1203
Weighting, Causal Effects, and Additive Interaction
data controlboot;
set boot;
if y¼0;
run;
proc logistic data¼controlboot descending outest¼params1; by sample;
model A¼ c1 c2 c3;
run;
proc logistic data¼boot descending inest¼params1; by sample;
model A¼c1 c2 c3 / MAXITER¼0;
output out¼boot predicted¼pa;
run;
proc logistic data¼controlboot descending outest¼params2; by sample;
model B¼ c1 c2 c3;
run;
proc logistic data¼boot descending inest¼params2; by sample;
model B¼ A c1 c2 c3 / MAXITER¼0;
output out¼boot predicted¼pb;
run;
data boot;
set boot;
if A¼1 then wa¼1/pa;
if A¼0 then wa¼1/(1-pa);
if B¼1 then wb¼1/pb;
if B¼0 then wb¼1/(1-pb);
ipw¼wa*wb;
run;
proc nlmixed data¼boot; by sample;
odds¼exp(b0)*(1þb1*Aþb2*Bþb3*A*B);
model y ~ binary(odds/(1þodds));
replicate ipw;
predict b3 out¼myout;
run;
proc univariate data¼myout;
var pred;
output out¼cis pctlpts¼2.5 97.5 pctlpre¼cis;
proc print data¼cis noobs label;
run;
Am J Epidemiol. 2011;174(10):1197–1203
1203