American Journal of Epidemiology ª The Author 2011. Published by Oxford University Press on behalf of the Johns Hopkins Bloomberg School of Public Health. All rights reserved. For permissions, please e-mail: [email protected]. Vol. 174, No. 10 DOI: 10.1093/aje/kwr334 Advance Access publication: October 19, 2011 Practice of Epidemiology A Weighting Approach to Causal Effects and Additive Interaction in Case-Control Studies: Marginal Structural Linear Odds Models Tyler J. VanderWeele* and Stijn Vansteelandt * Correspondence to Dr. Tyler J. VanderWeele, Departments of Epidemiology and Biostatistics, Harvard School of Public Health, 677 Huntington Avenue, Boston, MA 02115 (e-mail: [email protected]). Initially submitted October 28, 2010; accepted for publication June 13, 2011. Estimates of additive interaction from case-control data are often obtained by logistic regression; such models can also be used to adjust for covariates. This approach to estimating additive interaction has come under some criticism because of possible misspecification of the logistic model: If the underlying model is linear, the logistic model will be misspecified. The authors propose an inverse probability of treatment weighting approach to causal effects and additive interaction in case-control studies. Under the assumption of no unmeasured confounding, the approach amounts to fitting a marginal structural linear odds model. The approach allows for the estimation of measures of additive interaction between dichotomous exposures, such as the relative excess risk due to interaction, using case-control data without having to rely on modeling assumptions for the outcome conditional on the exposures and covariates. Rather than using conditional models for the outcome, models are instead specified for the exposures conditional on the covariates. The approach is illustrated by assessing additive interaction between genetic and environmental factors using data from a case-control study. case-control studies; interaction; linear model; structural model; synergism; weighting Abbreviations: OR, odds ratio; RERI, relative excessive risk due to interaction; RR, risk ratio. In this paper, we consider the use of inverse probability weighting to estimate causal effects in unmatched case-control studies. The approach we take effectively amounts to fitting what may be defined as a marginal structural linear odds model to case-control data. The approach is quite general. However, the methodological development here was motivated by the problem of attempting to assess interaction on the additive scale by using data from a case-control study. Additive interaction is often assessed by estimating a quantity sometimes referred to as the ‘‘relative excess risk due to interaction’’ (RERI) (1). If there are 2 dichotomous factors (call them A and B) and we let RRij denote the risk ratio (RR) comparing A ¼ i and B ¼ j with A ¼ B ¼ 0, then the relative excess risk due to interaction for the risk ratio is defined by RR11 – RR10 – RR01 þ 1. In a case-control study with ‘‘cumulative design’’ where controls are sampled from among those disease free at the end of follow-up, the odds ratio will generally be used to estimate the effect of the factors A and B. If the outcome is rare, then the risk ratios can be approximated by odds ratios (ORs). If we let ORij denote the odds ratio comparing A ¼ i and B ¼ j with A ¼ B ¼ 0, then the relative excess risk due to interaction for the odds ratio is defined by RERI ¼ OR11 OR10 OR01 þ 1: Throughout this paper, we will be using RERI on the odds ratio scale. The RERI can be used to give a measure of interaction on the additive scale for case-control data when the outcome is rare. Measures of interaction on the additive scale, such as RERI, are generally what is thought to be of most importance in considering public health implications (2–5). This is because additive interaction allows one to assess whether an intervention would have a larger absolute effect in one subpopulation versus another. Estimates of RERI are also useful in detecting synergism between 2 factors within the sufficient cause framework (5, 6). Under the assumption that both exposures 1197 Am J Epidemiol. 2011;174(10):1197–1203 1198 VanderWeele and Vansteelandt have neutral or causative effects for all individuals (i.e., the effects are positive monotonic), RERI > 0 implies such synergism (5–7); without such monotonicity assumptions, one can still test for synergism by testing RERI > 1 (6–8). Several articles (9–13) have considered the problem of estimating and providing confidence intervals for RERI. Earlier work by Hosmer and Lemeshow (9) and Assmann et al. (10) considered logistic regression models and the delta method or bootstrapping for confidence intervals for RERI. More recent work, such as that by Richardson and Kaufman (12), has considered using a linear odds model (5, 12–14), namely, odds ¼ exp(b0)(1 þ b1A þ b2B þ b3AB), to obtain estimates and confidence intervals for RERI. Under the linear odds model, the coefficient and confidence interval for b3 can be interpreted as RERI and its confidence interval (12). Both the approach using logistic regression and the approach using the linear odds model can be used to give confidence intervals even when controlling for other covariates, provided that the regression models are correctly specified. When there are covariates for which adjustments are to be made, these can likewise be included in either a logistic regression or linear odds model. In these cases, however, the relation between the covariates and the outcome needs to be correctly specified; failure to do so can lead to invalid inferences (14, 15). Here, we provide an alternative approach to estimating the RERI using weighting that will be applicable irrespective of whether the true underlying model relating the outcome and the covariates is a logistic model, a linear odds model, or some other model. This weighting approach will require that models for the exposures are correctly specified but will not require correctly specifying a model for the outcome. A WEIGHTING APPROACH TO RERI Suppose that data come from an unmatched case-control study (we make a few remarks about how the approach might be adapted to matched case-control designs in the Web Appendix, which is posted on the Journal’s Web site (http:// aje.oupjournals.org/)), where controls are sampled from those who are disease free at the end of follow-up. Suppose also that there are covariates for which control is to be made. Finally, we suppose that the outcome is rare; this assumption will be needed to estimate the weights and so that the relative excess risk due to interaction estimate for the odds ratio can genuinely be interpreted as a measure of additive interaction. We show in the Web Appendix that, instead of including covariates in a linear odds model or a logistic model, one can use a weighting approach for covariate adjustment as follows. One first estimates inverse probability of treatment weights using data just on the controls. This could be done by using 2 logistic regressions between the controls: 1) a logistic regression of the first exposure on the covariates and 2) a logistic regression of the second exposure on the covariates and the first exposure. Again, both regressions use data only on the controls. For each individual, a weight is obtained for the first exposure (call it wA) by taking the inverse of the predicted probability from the first logistic regression that the individual had the exposure level of A that was in fact present. Likewise, a weight is obtained for the second exposure (call it wB) by taking the inverse of the predicted probability from the second logistic regression that the individual had the exposure level of B that was in fact present. These are referred to as inverse-probability-of-treatment weights (16). Multiplying these 2 weights together (wA 3 wB) gives the overall weight for the individual. Note that, although the logistic regression models are fit for the controls only, the predicted probabilities and weights are calculated for each individual in the sample (both cases and controls). If a linear odds model conditional on the 2 exposures, Odds ¼ expðb0 Þð1 þ b1 A þ b2 B þ b3 ABÞ; is fit to case-control data with these weights, then, under a rare outcome assumption, the coefficient and confidence interval for b3 in this weighted linear odds regression will give an estimate of relative excess risk due to interaction for the standardized odds ratios adjusted for the covariates. Further detail is given in Appendix 1 and justification in the Web Appendix. Even though the procedure above requires estimation of the weights, if robust standard errors are used, it will still yield confidence intervals and estimates of the standard error that are conservative, as in other weighting approaches (16). As discussed in Appendix 1, the procedure of using the controls to calculate the weights and then fitting weighted linear odds models is applicable to the estimation of causal effects more generally and not simply to the RERI. Under the rare outcome assumption and provided that the set of covariates for which adjustment is made suffice to control for confounding, the procedure essentially amounts to fitting what can be defined as a ‘‘marginal structural linear odds model.’’ The approach also extends discussion of marginal structural models for interaction in cohort studies (17) to case-control studies. SAS implementation is also given in Appendix 2. ILLUSTRATION The approach is illustrated with data from a case-control study of lung cancer at Massachusetts General Hospital (18) of 1,836 cases and 1,452 controls. Eligible cases included any person over the age of 18 years, with a diagnosis of primary lung cancer that was further confirmed by a lung pathologist. The controls were recruited from among the friends or spouses of cancer patients or the friends or spouses of other surgery patients in the same hospital. Potential controls that carried a previous diagnosis of any cancer (other than nonmelanoma skin cancer) were excluded from participation. The study included information on smoking and genotype information on locus 15q25.1. Genetic variants on 15q25.1 have been found to be associated with both smoking and lung cancer (19–21). We examine whether there is interaction on the additive scale between the effects of smoking (ever vs. never) and the genetic variant (0 vs. 1/2 T alleles at rs8034191). Covariate data include age (continuous), gender, and educational history (college degree Am J Epidemiol. 2011;174(10):1197–1203 Weighting, Causal Effects, and Additive Interaction or more, yes/no). Analyses were limited to Caucasians. Using the procedure above gives an estimate of RERI ¼ 2.71 (percentile bootstrap confidence interval: 1.68, 4.04) with 2,000 bootstrapped samples. The estimate suggests positive interaction on the additive scale between smoking and the genetic variant. The estimate and confidence interval suggest synergism in the sufficient cause sense (5–8) even without assumptions about monotonicity. SIMULATION Data were simulated by using the sample sizes and datagenerating mechanism observed in the aforementioned casecontrol study of lung cancer. Specifically, in each of 1,000 simulation experiments, a large number of binary measurements for gender and educational history were drawn from their empirical multinomial distribution, and normally distributed age measurements were drawn conditionally on those. Smoking status and genotype were generated under logistic regression models, allowing for a dependence between both exposures, conditional on covariates. Lung cancer status was also generated from a logistic model, allowing for main effects and an interaction between both exposures, as well as the main effects of all 3 covariates. Subsequently, data were retained for 1,836 cases and 1,452 controls. The code generating the simulated data is given in the Web Appendix. Five simulation experiments were constructed, corresponding to 5 different population outcome means. Data were analyzed by 1) the proposed approach using logistic regression models for both exposures and also by 2) a conditional linear odds regression: OddsðYÞ ¼ expðc0 þ c4 CÞð1 þ c1 a þ c2 b þ c3 abÞ: Note that this model is deliberately misspecified under our data-generating model, in view of the aforementioned concerns about correctly specifying conditional linear odds regression models. Note further also that the weighting approach proposed in this paper is subject to model misspecification bias because it involves equating a model for both exposures correctly specified for the population with the corresponding model in controls. Reported coverage is for 95% Wald confidence intervals based on conservative standard errors for 1199 the proposed procedure and 95% likelihood intervals (12) for the conditional linear odds approach. The results are summarized in Table 1. They confirm the adequate performance of the proposed approach at disease prevalences below 10%. As predicted by the theory, the confidence intervals based on the proposed approach are conservative (this could be addressed through the use of bootstrap confidence intervals). The performance of both approaches deteriorates with larger disease prevalences, because the aforementioned degree of model misspecification grows with increasing prevalence. We note that Månsson et al. (22) also evaluated an inverseprobability weighting approach to total causal effects in the context of propensity score analysis using simulations and found reasonably good coverage probabilities for the approach under some, but not all, scenarios. Unfortunately, they did not directly report the outcome prevalence for their simulations, so it is difficult to compare their simulation scenarios with those reported here. DISCUSSION AND IMPLICATIONS FOR MODEL SPECIFICATION Motivation for this weighting approach to testing and estimation for additive interaction comes in part from a paper by Skrondal (14). Skrondal cautioned against the use of RERI and logistic regression to examine interaction on the additive scale in case-control studies. He noted that, if the underlying linear risk model with covariates was correctly specified, the additive interaction then took on a single value in all strata of the covariates but RERI would vary across strata (what he called the ‘‘uniqueness problem’’). Although this is true, it is nevertheless still the case that, if the linear risk model with covariates is correctly specified, RERI will be of the same sign in all strata of the covariates, and thus RERI > 0 (the condition for synergism under monotonicity) will either hold in all strata of covariates or in no stratum. The condition RERI > 1 (i.e., the condition for synergism without monotonicity) could vary across strata, but it is also the case that without monotonicity, synergism may be present in some strata of the covariates but not in others, even if the underlying linear risk model with covariates is correct (6, 7). Table 1. Results for Simulation Comparing Marginal Structural Linear Odds Models and Conditional Linear Odds Models at Different Outcome Prevalences When Both Are Potentially Subject to Misspecification Marginal Structural Linear Odds Model E (Y ),% Empirical SE Estimated SE Conditional Linear Odds Model Bias Empirical SE Coverage Probability, % 2.55 0.069 0.54 94.1 2.50 0.12 0.53 94.8 2.38 0.17 0.51 94.5 98.1 2.09 0.49 0.50 85.0 31.7 1.54 1.05 0.52 41.3 Coverage Probability, % RERI Bias 0.5 2.53 0.061 0.56 0.71 96.1 1.3 2.46 0.14 0.55 0.71 95.7 3.3 2.29 0.23 0.52 0.65 98.4 8.3 2.29 0.25 0.52 0.62 18.3 1.35 1.20 0.55 0.55 RERI Abbreviations: E(Y ), outcome prevalence; RERI, relative excess risk due to interaction; SE, standard error. Am J Epidemiol. 2011;174(10):1197–1203 1200 VanderWeele and Vansteelandt Perhaps more importantly, Skrondal noted that, if the linear risk model with covariates was correctly specified, the logistic model with covariates would not be correctly specified (what he called the ‘‘misspecification problem’’). Greenland (23) further remarked that parsimonious logistic regression models impose nonadditivity. By estimating RERI with such models, one thus risks introducing a bias toward nonadditivity on the additive scale. Skrondal noted that the linear odds model with covariates could be used to identify the parameters of the linear risk model up to a constant of proportionality (although this, in fact, only holds under an assumption that the outcome is rare). It is, however, conversely also the case that, if in fact the logistic regression model with covariates is correctly specified, then the linear risk model and the linear odds model will not be. The advantage of using the weighting approach we described above is that it will give valid estimates of the relative excess risk due to interaction for the standardized odds ratios irrespective of whether the true underlying model is a linear risk model with covariates, a logistic model with covariates, or some other model, provided that the models used for the weights are correctly specified. The weighting approach described above does not require a correctly specified conditional model for the outcome given the exposures and covariates, but it does require a correctly specified model for the weights; it thus effectively transfers the problem of correctly specifying a model from the outcome to the exposures, as is also the case with marginal structural models in cohort studies (16, 17). This weighting approach that avoids having to correctly specify an outcome model conditional on the covariates is furthermore desirable within the context of testing for synergism because an outcome model imposes certain assumptions within the sufficient cause framework, whereas a model for the exposures does not (17). In related work, we are developing a doubly robust approach that will yield valid inferences for unmatched case-control data if either a conditional model for the exposures or a conditional model for the outcome is correctly specified. The use of Bayesian approaches may also be of interest (24). ACKNOWLEDGMENTS Author affiliations: Department of Epidemiology, Harvard School of Public Health, Boston, Massachusetts (Tyler J. VanderWeele); Department of Biostatistics, Harvard School of Public Health, Boston, Massachusetts (Tyler J. VanderWeele); and Department of Applied Mathematics and Computer Science, Ghent University, Ghent, Belgium (Stijn Vansteelandt). The research was supported by grants ES017876 and HD060696 from the National Institutes of Health. Conflict of interest: none declared. REFERENCES 1. Rothman KJ. Modern Epidemiology. 1st ed. Boston, MA: Little, Brown and Company; 1986. 2. Blot WJ, Day NE. Synergism and interaction: are they equivalent? Am J Epidemiol. 1979;100(1):99–100. 3. Rothman KJ, Greenland S, Walker AM. Concepts of interaction. Am J Epidemiol. 1980;112(4):467–470. 4. Saracci R. Interaction and synergism. Am J Epidemiol. 1980; 112(4):465–466. 5. Rothman KJ, Greenland S, Lash TL. Modern Epidemiology. 3rd ed. Philadelphia, PA: Lippincott Williams & Wilkins; 2008. 6. VanderWeele TJ, Robins JM. The identification of synergism in the sufficient-component-cause framework. Epidemiology. 2007;18(3):329–339. 7. VanderWeele TJ. Sufficient cause interactions and statistical interactions. Epidemiology. 2009;20(1):6–13. 8. VanderWeele TJ. Empirical tests for compositional epistasis [letter]. Nat Rev Genet. 2010;11(2):166. 9. Hosmer DW, Lemeshow S. Confidence interval estimation of interaction. Epidemiology. 1992;3(5):452–456. 10. Assmann SF, Hosmer DW, Lemeshow S, et al. Confidence intervals for measures of interaction. Epidemiology. 1996; 7(3):286–290. 11. Nie L, Chu H, Li F, et al. Relative excess risk due to interaction: resampling-based confidence intervals. Epidemiology. 2010;21(4):552–556. 12. Richardson DB, Kaufman JS. Estimation of the relative excess risk due to interaction and associated confidence bounds. Am J Epidemiol. 2009;169(6):756–760. 13. Kuss O, Schmidt-Pokrzywniak A, Stang A. Confidence intervals for the interaction contrast ratio. Epidemiology. 2010; 21(2):273–274. 14. Skrondal A. Interaction as departure from additivity in casecontrol studies: a cautionary note. Am J Epidemiol. 2003; 158(3):251–258. 15. Greenland S. Additive risk versus additive relative risk models. Epidemiology. 1993;4(1):32–36. 16. Robins JM, Hernán MA, Brumback B. Marginal structural models and causal inference in epidemiology. Epidemiology. 2000;11(5):550–560. 17. VanderWeele TJ, Vansteelandt S, Robins JM. Marginal structural models for sufficient cause interactions. Am J Epidemiol. 2010;171(4):506–514. 18. Miller DP, Liu G, De Vivo I, et al. Combinations of the variant genotypes of GSTP1, GSTM1, and p53 are associated with an increased lung cancer risk. Cancer Res. 2002;62(10): 2819–2823. 19. Hung RJ, McKay JD, Gaborieau V, et al. A susceptibility locus for lung cancer maps to nicotinic acetylcholine receptor subunit genes on 15q25. Nature. 2008;452(1787): 633–637. 20. Amos CI, Wu X, Broderick P, et al. Genome-wide association scan of tag SNPs identifies a susceptibility locus for lung cancer at 15q25.1. Nat Genet. 2008;40(5):616–622. 21. Thorgeirsson TE, Geller F, Sulem P, et al. A variant associated with nicotine dependence, lung cancer and peripheral arterial disease. Nature. 2008;452(7187): 638–642. 22. Månsson R, Joffe MM, Sun W, et al. On the estimation and use of propensity scores in case-control and case-cohort studies. Am J Epidemiol. 2007;166(3):332–339. 23. Greenland S. Interactions in epidemiology: relevance, identification, and estimation. Epidemiology. 2009;20(1): 14–17. 24. Chu H, Nie L, Cole SR. Estimating the relative excess risk due to interaction: a Bayesian approach. Epidemiology. 2011; 22(2):242–248. Am J Epidemiol. 2011;174(10):1197–1203 Weighting, Causal Effects, and Additive Interaction APPENDIX 1 Marginal Structural Linear Odds Models Let Y denote the outcome of interest, and let T denote the exposure(s) of interest. The variable T may include more than one exposure. In the application to RERI, T consists of 2 exposures (A, B). We let Yt denote the counterfactual value of Y if T had been set to level t. The odds for Yt can then be defined as OddsðYt Þ ¼ PðYt ¼ 1Þ=f1 PðYt ¼ 1Þg: 1201 OddsðYða;bÞ oddsðYð0;0Þ ¼ ðh0 þ h1 a þ h2 b þ h3 ab h0 ¼ ðkc0 þ kc1 a þ kc2 b þ kc3 ab kc0 ¼ 1 þ ðc1 c0 Þa þ ðc2 c0 Þb þ ðc3 c0 Þab: If, instead, we parameterized the marginal structural linear odds model as OddsðYða;bÞ Þ ¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ A marginal structural linear odds model takes the form: and fit a weighted conditional linear odds regression, OddsðYt Þ ¼ gðtÞ#h; OddsðYÞ ¼ expðc0 Þð1 þ c1 a þ c2 b þ c3 abÞ; where g(t) is a vector-valued function of t, and h is a vector of parameters. For example, for 2 exposures A and B, we would have that T ¼ (A, B), g(t) ¼ g(a, b) ¼ (1, a, b, ab)#, and h ¼ (h0, h1, h2, h3)# so that the marginal structural linear odds model is then the parameter estimates, (c1, c2, c3), from the reparameterized weighted linear odds regression will directly give consistent estimates of the causal odds ratio parameters (b1, b2, b3). Note that, to obtain consistent estimates of the causal odds ratios, both assumption 1 of rare outcome and assumption 2 of no confounding conditional on C were required. If only assumption 1 holds (i.e., if assumption 2 does not hold so there is still confounding), the weighting procedure above will still give consistent estimates for the standardized odds ratios adjusted for C, namely: P PðY ¼ 1j A ¼ a; B ¼ b; C ¼ cÞPðC ¼ cÞ c P 1 PðY ¼ 1j A ¼ a; B ¼ b; C ¼ cÞPðC ¼ cÞ OddsðYða;bÞ Þ ¼ h0 þ h1 a þ h2 b þ h3 ab: The parameters of the marginal structural linear odds model can be fit with case-control data by using inverse probability of treatment weighting if 1) a rare outcome assumption holds and 2) the set of covariates C suffices to control for confounding for the effects of T on Y (in counterfactual notation, this is Yt is independent of T conditional on C). If so, inverse probability of treatment weights for individual i is given by wi ¼ 1/P(T ¼ tijC ¼ ci, Y ¼ 0), where the denominator probability can be estimated by regression models (multiple regression models if T is multivariate). For example, for 2 exposures, the weights would be as follows: wi ¼ 1 1 ; 3 PðA ¼ ai jC ¼ ci Þ PðB ¼ bi jA ¼ ai ; C ¼ ci Þ where each of the 2 denominator probabilities could be estimated by logistic regression. Under assumptions 1) and 2) fitting a conditional linear odds regression, odds(Y) ¼ g(t)# c with the observed case-control data, where each subject is weighted by wi, will give consistent estimates of the corresponding parameters, h, of the marginal structural model up to a constant of proportionality, k, so that kc ¼ h. For example, with 2 exposures, A and B, we have that fitting the conditional linear odds regression: OddsðYÞ ¼ c0 þ c1 a þ c2 b þ c3 ab; where each subject weighted by wi will give consistent estimates of the corresponding parameters of the marginal structural model (h0, h1, h2, h3) up to a constant of proportionality, that is (kc0, kc1, kc2, kc3) ¼ (h0, h1, h2, h3). Note that, from the estimates of (c0, c1, c2, c3), we can estimate the causal odds ratios: Am J Epidemiol. 2011;174(10):1197–1203 c ¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ: In the case of estimating RERI with 2 dichotomous factors, the model specifications OddsðYða;bÞ Þ ¼ h0 þ h1 a þ h2 b þ h3 ab and OddsðYða;bÞ Þ ¼ expðb0 Þð1 þ b1 a þ b2 b þ b3 abÞ are both saturated. Model misspecification is not an issue. In the more general case, the marginal structural model, odds(Yt) ¼ g(t)#h, may not be saturated. In this case, the marginal structural model must be correctly specified to get valid estimates of causal effects. Moreover, when the marginal structural model is not saturated, more efficient estimates of the coefficients can sometimes be obtained by using what are sometimes called ‘‘stabilized weights’’ (16) rather than the weights wi given above. Using stabilized weights, one would weight the conditional linear odds regression by swi ¼ PðT ¼ ti jY ¼ 0Þ PðT ¼ ti jC ¼ ci ; Y ¼ 0Þ rather than by wi ¼ 1 PðT ¼ ti jC ¼ ci ; Y ¼ 0Þ: 1202 VanderWeele and Vansteelandt APPENDIX 2 SAS Implementation for a Weighting Approach to RERI We describe how the weighting approach to RERI given above for an unmatched case-control design can be implemented in SAS statistical software (SAS Institute, Inc., Cary, North Carolina). Suppose the data are stored in a data set defined as ‘‘mydata’’ and the outcome variable is ‘‘y,’’ the 2 binary exposures are ‘‘A’’ and ‘‘B’’ and the covariates are c1, c2, c3. The following SAS code will create the predicted probabilities above (pa and pb), the 2 weights (wa and wb), the overall inverse probability weight (ipw), and the estimate of the RERI (b3). data controldata; set mydata; if y¼0; run; proc logistic data¼controldata descending outest¼params1; model A¼ c1 c2 c3; run; proc logistic data¼mydata descending inest¼params1; model A¼ c1 c2 c3/ MAXITER¼0; output out¼mydata predicted¼pa; run; proc logistic data¼controldata descending outest¼params2; model B ¼ A c1 c2 c3; run; proc logistic data¼mydata descending inest¼params2; model B ¼ A c1 c2 c3/ MAXITER¼0; output out¼mydata predicted¼pb; run; data mydata; set mydata; if A¼1 then wa¼1/pa; if A¼0 then wa¼1/(1-pa); if B¼1 then wb¼1/pb; if B¼0 then wb¼1/(1-pb); ipw¼wa*wb; run; proc nlmixed data¼mydata; odds¼exp(b0)*(1þb1*Aþb2*Bþb3*A*B); model y ~ binary(odds/(1þodds)); replicate ipw; run; The final procedure of this code (proc nlmixed) is essentially that given by Richardson and Kaufman (12) plus an additional replicate statement to weight the observations. The coefficient b3 produced by SAS will be an estimate of the RERI. Unfortunately, SAS proc nlmixed does not have an option to produce robust standard errors that are needed for a valid confidence interval in this case. Here, we provide additional SAS code to give bootstrapped standard errors for the estimate obtained from the SAS code above. In the Web Appendix, we also give alternative SAS code using the delta method to give robust standard errors. The following SAS code can be used to obtain percentile bootstrap confidence intervals for this estimate. The code uses 2,000 samples, but users could increase this to a larger number as well by changing the second line of the code. data boot; do sample¼1 to 2000; do i¼1 to nobs; pt¼round(ranuni(12)*nobs); set mydata nobs¼nobs point¼pt ; output; end; end; stop; run; Am J Epidemiol. 2011;174(10):1197–1203 Weighting, Causal Effects, and Additive Interaction data controlboot; set boot; if y¼0; run; proc logistic data¼controlboot descending outest¼params1; by sample; model A¼ c1 c2 c3; run; proc logistic data¼boot descending inest¼params1; by sample; model A¼c1 c2 c3 / MAXITER¼0; output out¼boot predicted¼pa; run; proc logistic data¼controlboot descending outest¼params2; by sample; model B¼ c1 c2 c3; run; proc logistic data¼boot descending inest¼params2; by sample; model B¼ A c1 c2 c3 / MAXITER¼0; output out¼boot predicted¼pb; run; data boot; set boot; if A¼1 then wa¼1/pa; if A¼0 then wa¼1/(1-pa); if B¼1 then wb¼1/pb; if B¼0 then wb¼1/(1-pb); ipw¼wa*wb; run; proc nlmixed data¼boot; by sample; odds¼exp(b0)*(1þb1*Aþb2*Bþb3*A*B); model y ~ binary(odds/(1þodds)); replicate ipw; predict b3 out¼myout; run; proc univariate data¼myout; var pred; output out¼cis pctlpts¼2.5 97.5 pctlpre¼cis; proc print data¼cis noobs label; run; Am J Epidemiol. 2011;174(10):1197–1203 1203
© Copyright 2026 Paperzz