Estimating marginal and incremental effects on

Biostatistics (2005), 6, 1, pp. 93–109
doi: 10.1093/biostatistics/kxh020
Estimating marginal and incremental effects on
health outcomes using flexible link and variance
function models
ANIRBAN BASU∗
Section of General Internal Medicine, Department of Medicine, University of Chicago, 5841 South
Maryland Ave—MC 2007, Chicago IL 60637, USA
[email protected]
PAUL J. RATHOUZ
Department of Health Studies, University of Chicago, 5841 South Maryland Ave—MC 2007,
Chicago IL 60637, USA
[email protected]
S UMMARY
We propose an extension to the estimating equations in generalized linear models to estimate parameters
in the link function and variance structure simultaneously with regression coefficients. Rather than
focusing on the regression coefficients, the purpose of these models is inference about the mean of
the outcome as a function of a set of covariates, and various functionals of the mean function used to
measure the effects of the covariates. A commonly used functional in econometrics, referred to as the
marginal effect, is the partial derivative of the mean function with respect to any covariate, averaged
over the empirical distribution of covariates in the model. We define an analogous parameter for discrete
covariates. The proposed estimation method not only helps to identify an appropriate link function and
to suggest an underlying distribution for a specific application but also serves as a robust estimator when
no specific distribution for the outcome measure can be identified. Using Monte Carlo simulations, we
show that the resulting parameter estimators are consistent. The method is illustrated with an analysis of
inpatient expenditure data from a study of hospitalists.
Keywords: Econometric models; Estimating equations; Generalized linear models; Link function; Variance function.
1. I NTRODUCTION
Many analysis problems in health economics and biostatistics involve modeling a response variable
Y as a function of a vector X = (X 1 , X 2 , . . . , X p )T of covariates in a regression model for the mean
function µ(x) ≡ E(Y |X = x). Usually, interest lies in estimating the effect of one or more of the
covariates X j on Y . In many cases, this effect is measured by regression coefficients β j , but in others
it is quantified via more general functionals of µ(x). One such functional, commonly used in health
economics, is the partial derivative of µ(x) with respect to covariate x j in vector x = (x1 , x2 , . . . , x p )T .
Denoted by D j (µ; x) ≡ ∂µ(x)/∂ x j , this parameter is the rate of change in µ(X ) with respect to X j
∗ To whom correspondence should be addressed.
c Oxford University Press 2005; all rights reserved.
Biostatistics Vol. 6 No. 1 94
A. BASU AND P. J. R ATHOUZ
evaluated at X = x. When X j is an indicator variable, we define D j (µ; x− j ) to be the difference in µ(x)
at the two levels of X j , i.e. D j (µ; x− j ) ≡ µ(x j = 1, x− j ) − µ(x j = 0, x− j ), where x− j is the vector x
without x j .
In many applications, any particular combination of values, such as X = x, represents only a
negligible fraction of the population, and researchers would like to assess the effect of X j in the whole
population. To this end, econometricians commonly focus on the expected value of D j (µ; X ) over the
population distribution of X . This parameter is termed the marginal effect of the covariate, in the sense
used by econometricians (Greene, 2000, p. 824) and is given by ξ j ≡ E X {D j (µ; X )}; it is the population
average rate of change in µ(X ) with respect to X j , controlling for other
factors X − j . When X j is binary,
the parameter of interest is the incremental effect given by π j ≡ E X − j D j (µ; X − j ) , where the expected
value is over X − j , marginally with respect to X j . The parameter π j is the population average contrast in
the mean of Y for X j = 1 and X j = 0. Here, the expectation is again taken over X , but as X j is fixed
at 0 or 1 in D j (µ; X − j ), π j only involves the marginal distribution of X − j . The interpretation of both ξ j
and π j are as effects of X j on the mean of Y , adjusting for all other covariates in the model, where this
adjustment is to the population distribution of X . When µ(x) is linear in x j , then either ξ j or π j is simply
equal to β j .
A recent example from health economics arises in a two-year study of hospitalists at the University of
Chicago (Meltzer et al., 2002). Hospitalists are physicians who spend three months a year attending on
inpatient wards, rather than the one month typical of most physicians in academic medical centers. The
policy issues are whether hospitalists provide less expensive care than the traditional arrangement, and
if so, what the magnitudes of these effects are. Preliminary evidence shows that, at the beginning of the
study, there were no differences in utilization (inpatient expenditures and length of stay) between patients
treated by hospitalists and those treated by non-hospitalists, suggesting that there were no significant or
appreciable differences in baseline skills or experience between the hospitalist and traditional attending
teams. Instead, it appears that the differences evolve over time and are directly related to accumulated
physician experience on the date of admission of the observation. At the end of the two year period
and after adjusting for patient demographics and clinical conditions, hospitalist patients had significantly
lower utilization rates than non-hospitalist patients (Meltzer et al., 2002). The behavioral question is
whether this difference is due to the higher cumulative inpatient experience (i.e. the number of prior
disease-specific cases treated) of attending hospitalists over time. That is, as the number of cases treated
increases, do expenditures fall on average across the population and if so, what is the adjusted hospitalist
effect in terms of dollars saved? Also, does the introduction of a covariate for disease-specific cumulative
experience eliminate the hospitalist effect? Here, policy interest is on additive effects of experience on
mean expenditures. Letting X 1 be the hospitalist indicator variable and X 2 the disease-specific experience
of the attending physician, the goal is to estimate the incremental effect π1 of being a hospitalist and
the marginal effect ξ2 of disease-specific experience on inpatient expenditures on average across the
patient and attending populations, adjusting for patient demographics, clinical conditions and type of care
received. The scientific interpretation of π1 is the average cost-saving in dollars generated by physicians
who are hospitalist as compared to non-hospitalist treating an otherwise identical population of patients.
Similarly, the scientific interpretation of ξ2 is the cost-saving in dollars generated by an increase in one
unit of disease-specific experience per physician, averaged over all patients and physicians.
When correct, a linear model µ(x) = x T β estimated with ordinary least squares (OLS) will yield
regression coefficients β̂ j which are consistent estimators of ξ j and π j . However, many outcome variables
in health economics and biostatistics are characterized by non-negative values, heteroscedasticity, heavy
skewness in the right tail, and kurtotic distributions, rendering OLS on the raw scale of Y inapplicable.
Even when the linear model is correct, resulting estimators could be unstable, due to skewness and
kurtosis, and/or inefficient due to heteroscedasticity. Examples include inpatient length of stay and
expenditure, income and earnings, and counts of psychiatric or other symptoms. For example, the OLS
95
–4
Logscale Residuals
–2
0
2
4
Estimating marginal and incremental effects on health outcomes
7
8
9
10
Linear prediction
11
12
Fig. 1. Heteroscedasticity in log-scale OLS residuals.
residuals for total inpatient expenditure (in $) from the hospitalist data has a coefficient of skewness of
5.4 and a coefficient of kurtosis of 68.
Econometricians have historically relied on logarithmic or other transformations of Y , followed by
regression of the transformed Y on X using OLS, to overcome problems of heteroscedasticity, severe
skewness, and kurtosis (Box and Cox, 1964). In the hospitalist data, for example, the log-scale residuals
for total inpatient expenditure is better behaved than the raw scale, with coefficients of skewness and
kurtosis of 0.09 and 4. The main drawback of transforming Y is that the analysis does not result in a model
for µ(x) in the original scale, a scale that in many applications is the scale of interest. In the hospitalist
study, the scale of interest when modeling inpatient expenditure is dollars while the scale of estimation
is log-dollars in a log-OLS model. In order to draw inferences about the mean µ(x) or any functional
thereof in the natural scale of Y , one has to implement a retransformation (Duan, 1983; Manning, 1998)
involving the errors terms in the scale of estimation. This retransformation is complicated in the presence
of heteroscedasticity on the scale of estimation (Manning, 1998; Mullahy, 1998). For example, Figure 1
illustrates the presence of heteroscedasticity in the log-scale residuals over the log-scale linear predictor in
the hospitalist data, where reduced log-scale variance is observed at higher values of the mean. In practice,
as we seldom know the true form of heteroscedasticity, any retransformation can potentially yield biased
estimators of µ(x) unless considerable effort is devoted to studying the specific form of heteroscedasticity.
To avoid such problems of retransformation, biostatisticians and some economists have focused on
the use of generalized linear models (GLMs) with quasi-likelihood estimation (Wedderburn, 1974).
In the GLM approach, a link function relates µ(x) to a linear specification x T β of covariates. The
retransformation problem is eliminated by transforming µ(x) instead of Y . Moreover, GLMs allow for
heteroscedasticity through a variance structure relating Var(Y |X = x) to the mean, correct specification of
which results in efficient estimators (Crowder, 1987) and may correspond to an underlying distribution of
the outcome measure. Although log link models with the gamma error distribution are the most common
GLM application in health economics (Blough et al., 1999; Manning and Mullahy, 2001; Basu et al.,
2002), this specification is not universally correct, and it is often difficult to identify the appropriate link
96
A. BASU AND P. J. R ATHOUZ
function and variance structure a priori (Blough et al., 1999; Manning and Mullahy, 2001). Economic
theory has a difficult enough time predicting the direction of effects, and provides almost no guidance
about the functional form of µ(x) or about distributional characteristics of Y given X . One approach to
this problem is to employ a series of diagnostic tests for candidate link and variance function models;
examples include the Pregibon link test (1980), the Hosmer–Lemeshow test (1995) and the modified Park
test (Manning and Mullahy, 2001). However, in many cases, even if these tests detect problems, they do
not provide any guidance on how to fix those problems. For example, Figure 3(a) shows that applying the
standard log-link model with gamma variance to the hospitalist data results in overestimating the mean
(i.e. negative residuals) at the right tails of the distribution of the linear predictor. An alternative approach,
which we pursue in this paper, is to estimate the link function and variance structure along with other
components of the model.
We propose a semi-parametric method to estimate the mean model µ(x) and the variance structure
for Y given X , concentrating on the case where Y is a positive random variable. We extend the traditional
GLM framework via a mean model that contains an additional parameter governing the link function
using the Box–Cox transformation (Box and Cox, 1964), and we propose parametric models for the
variance as a function of µ(x). We estimate the regression and link parameters via an extension of
quasi-likelihood (Wedderburn, 1974), and the variance parameters using additional estimating equations.
Finally, we show how to use this fitted model to make inferences about marginal (ξ j ) and incremental
(π j ) effects. The flexible estimation method we propose has three primary advantages: first, it helps to
identify an appropriate link function and suggests an underlying model for the error distribution for a
specific application; second, the proposed method itself is a robust estimator when no specific distribution
for the outcome measure can be identified. That is, our approach is semi-parametric in that, while we
employ parametric models for the mean and variance of (Y |X ), we do not employ further distributional
assumptions or full likelihood estimation methods. Finally, our method decouples the scale of estimation
for the mean model, determined by the link function, from the scale of interest for the scientifically
relevant effects: that is regardless of what link function is used, marginal and incremental effects on any
scale can be obtained. In this paper, we focus on additive effects.
The rest of this paper is structured as follows. The model definition, basic assumptions and estimation
method for µ(x), ξ j and π j are presented in Section 2. Section 3 presents a simulation study comparing
the performance of the proposed estimator with several other GLM estimators in terms of consistency in
estimating functionals of µ(x), specifically the marginal effects ξ j , and in terms of efficiency loss due
to estimation of additional parameters, versus cases when the appropriate link and variance are known
a priori. In Section 4, we illustrate the application of the proposed method with analysis of inpatient
expenditure data from the hospitalist study.
2. E XTENDED E STIMATING E QUATIONS (EEE) IN GENERALIZED LINEAR MODELS
2.1
Model
Consider N iid observations (Yi , X i ), where Yi is a positive response variable and X i = (X i1 , . . . , X i p )T
is a vector of covariates that may include an intercept. Interest is on modeling the mean function µ(x) ≡
E(Yi |X i = x) and functionals thereof. Letting µi = µ(X i ), we posit a GLM (Mccullagh and Nelder,
1989) wherein g(µi ) = ηi , ηi = X iT β and β is a p × 1 vector of regression parameters. Here, g(.)
is a strictly monotonic differentiable link function that relates µi to the linear predictor ηi . In addition,
the variance of the outcome variable is given by, Vi = Var(Yi |X i ) = ϕh(µi ), where ϕ is the dispersion
parameter and variance function h(µi ) is positive for all µi . In traditional GLMs, the link function and
variance structure are fixed by the investigator.
Our goal is to extend the GLM model for the mean and the variance of Yi to include families of
Estimating marginal and incremental effects on health outcomes
97
Table 1. Special cases of distributions under PV and QV Variance formulations
Variance formulation
PV
QV
θ1
θ2
θ1
θ2
1
1
1
0
Distribution
PDF
Poisson
e−µ µ y /y!
1
θ1
y
y
y
exp
−
µθ1
(1/θ1 ) µθ1
2
(y−µ)
1
√
exp
−
2yµ2 θ1
2π y 3 θ1
1
1
>0
2
0
>0
Gamma
>0
3
+
+
Inverse Gaussian
+
+
1
>0
Negative binomial
θ
y+ θ
2
θ1 (y+1)
2
1
θ2
θ2
µ+ θ1
2
PV = power variance, where V (yi ) = θ1 µi 2 .
QV = quadratic variance, where V (yi ) = θ1 µi + θ2 µi2 .
‘+’ = True value unknown since distribution does not conform to the particular variance
structure assumed.
link functions g(.; λ) and of variance functions h(.; θ1 , θ2 ). As such, define a parametric family of link
functions indexed by λ:
(µiλ − 1)/λ, if λ = 0
ηi = g(µi ; λ) =
(2.1)
log(µi ),
if λ = 0
(Mccullagh and Nelder, 1989, Chap. 2; Box and Cox, 1964). Notice that g(µi ; λ) is continuous in λ and
has continuous first derivatives in λ for all µi > 0 and for all λ, including λ = 0. As λ is allowed to
vary, the scale of the linear predictor ηi , relative to µi also varies. However, at ηi = 0, we have µi = 1
and ∂µi /∂ηi = 1 for all λ. Therefore, the family of link functions g(µi ; λ) is standardized such that
µi = 1 and ∂µi /∂ X i j = β j ( j = 1, . . . , p) when X iT β = 0, across all values of λ. For link function
g(µi ; λ) to be valid for a given λ, we restrict the linear predictor ηi such that (ηi λ + 1) > 0 which implies
ηi > (−1/λ) if λ 0, and ηi < (−1/λ) if λ < 0. This ensures that µi > 0 ∀i.
Similar to the link function, define a family h(µi ; θ1 , θ2 ) of variance functions indexed by (θ1 , θ2 ).
We consider two such families. The first, which we refer to as the power variance (PV) family, sets
h(µi ; θ1 , θ2 ) = θ1 µiθ2 . It includes as special cases the variances of several standard distributions used for
modeling health outcomes (see Table 1). An alternative is the quadratic variance (QV) family given by
h(µi ; θ1 , θ2 ) = θ1 µi + θ2 µi2 ; standard distributions corresponding to this form of variance are also listed
in Table 1. Note that we have subsumed the dispersion parameter ϕ into the variance functions (i.e. in PV
φ = θ1 and in QV, φ is a multiplier that is present both in θ1 and θ2 ).
2.2
Estimation
In traditional GLM (Mccullagh and Nelder, 1989), the regression parameters β are estimated using the
well-known quasi score equations (Wedderburn, 1974)
N
i=1
G iβ j =
N
(Yi − µi )Vi−1 (∂µi /∂β j ) = 0,
j = 1, . . . , p.
(2.2)
i=1
Mccullagh (1983) showed that solving equation (2.2) is equivalent to maximizing a quasi-likelihood
function that behaves in many ways as a likelihood for the regression parameters.
98
A. BASU AND P. J. R ATHOUZ
Building on (2.2), we define an extended set of estimating functions for parameter vector γ =
(β T , λ, θ1 , θ2 )T . Hence we name it the Extended Estimating Equations (EEE) model. For the ith
individual, define
G iβ j = (Yi − µi )Vi−1 (∂µi /∂β j )
j = 1, . . . , p
G iλ = (Yi − µi )Vi−1 (∂µi /∂λ)
G iθ1 = (Yi − µi )2 − Vi Vi−2 (∂ Vi /∂θ1 )
G iθ2 = (Yi − µi )2 − Vi Vi−2 (∂ Vi /∂θ2 ).
(2.3)
Quasi-score equation G iλ is similar to (2.2) since, like β, λ is also a mean model parameter. Estimating
functions G iθ1 and G iθ2 for variance parameters (θ1 , θ2 ) are unbiased and therefore provide consistent
estimators of θ1 and θ2 under the assumption that the mean model µi and the variance model h(µi ; θ1 , θ2 )
are correct (Hall and Sevrini, 1998). Define G iγ = {G iβ1 , G iβ2 , . . . , G iβ p , G iλ , G iθ1 , G iθ2 }T and the extended
N
estimating function for γ as G γ = i=1
G iγ .
p
We estimate γ by solving G γ = 0, yielding estimator γ̂ N . Under mild regularity conditions, γ̂ N −→γ0
as N → ∞ and (γ̂ N − γ0 ) is asymptotically normal with mean 0 and covariance matrix A N given by
AN =
−1
E(−∂G γ /∂γ )
N
−T
N
E(G iγ G iγT ) E(−∂G γ /∂γ )
.
(N − 1) i=1
(2.4)
A sketch of the proof is given in Appendix A. Replacing γ by γ̂ N and E(G iγ G iγT ) with G iγ G iγT in (2.4)
yields a sandwich estimator of the variance–covariance of γ̂ N (Huber, 1972; Liang and Zeger, 1986b).
Some studies yield clustered observations. For example in the hospitalist study, individual patients are
clustered within physicians. While estimator γ̂ N , albeit inefficient, is still consistent for γ because G γ is
unbiased, the variance–covariance estimator given by (2.4) is inconsistent. However, (2.4) can be modified
to yield a sandwich estimator that accounts for the clustering effect. Specifically, let M denote the total
number of clusters in the sample, where the mth cluster (m = 1, 2, . . . , M) contains Nm iid observations
(Ymi , X mi ), and yields estimating function G mi
γ , i = 1, 2, . . . , Nm . The asymptotic variance–covariance
of γ̂ N as N → ∞, M → ∞ is given by
AM =
−1
E(−∂G γ /∂γ )
M
(M − 1)
M
mT
E(Q m
γ Qγ )
−T
E(−∂G γ /∂γ )
,
(2.5)
m=1
Nm mi
m mT
m mT
where Q m
γ =
i=1 G γ . Replacing γ by γ̂ N and E(Q γ Q γ ) with Q γ Q γ in (2.5) yields the sandwich
estimator of the variance–covariance of (γ̂ N − γ0 ) accounting for clustering in the data.
A Fisher scoring algorithm used to solve G γ = 0 is described in Appendix B. The algorithm is
implemented in Stata SE (Statacorp., 2001) and is available on request from the corresponding author.
Further details are posted at http://home.uchicago.edu/~abasu/EEE/EEEWeb.pdf.
2.3
Estimation of marginal and incremental effects
For the model proposed in Section 2.1, the partial derivative of µ(x) with respect to a covariate x j is given
1−λ
. Thus, an estimator of the marginal effect ξ j of any continuous covariate X j on Y ,
by β j µ(x T β, λ)
Estimating marginal and incremental effects on health outcomes
99
denoted by ξ̂ j , is given by
ξ̂ j = Ê X {D j (µ̂; X )} = N −1
N
β̂ j {µ(X iT β̂, λ̂)}1−λ̂ .
(2.6)
i=1
Here the hat (∧ ) on µ̂ indicates that β and λ have been estimated and the hat on Ê X indicates that the
sample expected value has replaced the population expected value. To estimate the incremental effect π j
of an indicator variable X j , we use the method of recycled predictions (Statacorp., 2001). This method,
in obvious analogy to the definition of incremental effect π j for covariate X j , estimates π j as
π̂ j = Ê X − j {D j (µ̂; X − j )} = N −1
N
{µ̂(X i,− j , X i j = 1) − µ̂i (X i,− j , X i j = 0)}.
(2.7)
i=1
Variance estimators for the marginal and incremental effect estimators ξ̂ j and π̂ j are obtained using
Taylor series approximations. These depend both on the variance of (β̂, λ̂) and also on the variance of
covariates X in the population of interest. We show in Appendix C that the variance for (2.7) is given by
∂π ∂π j T
j
Var π̂ j = Var Ê X − j {D j (µ; X − j )} +
AN
.
(2.8)
∂γ α
∂γ α
In (2.8), the first term is the sample variance of π̂ j due to using the empirical expected value over
X − j , Ê X − j {D j (µ; X − j )}, rather than the population expected value. The second term is due to the fact
that γ is estimated. An estimator of the variance (2.8) is obtained by replacing γ with γ̂ , and replacing
the first term in (2.8) by
N
−1
−1
2
N
(N − 1)
(2.9)
{D j (µ̂; X − j )−π̂ j } .
i=1
Variance estimators for the estimated effect π̂ j that follow from (2.8) may be modified to account for
clustered observations. First, A N in (2.8) may be replaced with A M from (2.5). Second, the variance
estimator of π̂ in presence of clustering is given by
M
−2
−1
2
¯
M N (M − 1)
(2.10)
(π̂m −π̂ )
m=1
Nm
M
−1
Nm
where π̂m = i=1 π̂mi , π̂¯ = Nm N
m=1
i=1 π̂mi .
An estimator for Var(ξ̂ j ) analogous to (2.8) may be obtained through a similar approach.
3. S IMULATIONS
3.1
Design
We performed a simulation study to compare the performance of the EEE estimator to alternative
estimators under a variety of data generating processes (DGPs). We consider processes that differ in
their degrees of skewness and kurtosis, and, through different link functions, in their dependence on
a single covariate X . For all DGPs, X is uniformly distributed over [0, 1], η = β0 + β1 X, β1 = 1,
and β0 is selected such that E(Y ) = 1 marginally over X . The first DGP for (Y |X ) is the gamma
100
A. BASU AND P. J. R ATHOUZ
distribution with the shape parameter 0.5 and the scale parameter µ is related to η through three different
link functions: log (log(µ) = η), inverse (µ−1 = η) and square-root (µ0.5 = η). The second is the
inverse Gaussian distribution with shape parameter 1 and scale parameter µ related to η via the log link
and the identity link (µ = η). Finally, we considered (Y |X ) as log-normal with log-scale mean equal to
η and heteroscedastic log-scale variance equal to either v(X ) = (1 + X ) or to v(x) = (1 + X )2 , yielding
µ(x) = exp{β0 + β1 x + 0.5v(x)}. We also studied Poisson and negative binomial DGPs for (Y |X );
results were very similar to those for the gamma distribution and are available on request. For each DGP,
we generated 500 replicates each of sample size N = 2000 and N = 10 000.
For each replicate data set, we estimated µ(x) and Var(Y |X = x) under five different models: the
Gamma, Poisson and inverse Gaussian regression models of Y on X , each with a log link function, and
the flexible link model (2.1) with the PV and the QV variance functions, estimated using EEE. We also
studied the negative binomial regression model estimator, but omit the results due to similarity to those
of Poisson regression. Each of the five estimators sets η = β0 + β1 X . Although this specification is
incorrect for heteroscedastic log-normal data with quadratic variance, we wanted to see how the proposed
estimators perform under such misspecification.
For each fitted mean and variance model, we computed a variety of parameter estimates: (1) the fitted
mean µ̂(x) at the midpoint of each decile of the distribution of X ; (2) the fitted partial derivative Dx (µ̂; x)
at the midpoint of each decile of X and also at specific values of x; (3) the marginal effect ξ̂ computed
ˆ
using (2.6); and (4) the fitted variance Var(Y
|X = x) at the midpoint of each decile of the distribution of
X . By computing the mean of these statistics across all 500 replicates, we obtained the percent bias of each
estimator, which is presented graphically. In addition, to evaluate relative efficiency among estimators, we
computed the percent coefficient of variation of Dx (µ̂; x) (given by 100 × sd{Dx (µ̂; x)}/Dx (µ; x)) and
of ξ̂ over 500 replicates.
3.2
Results
For each data generating mechanism, Y is skewed to the right and heavy tailed, with the log-normal
distribution exhibiting the greatest skewness and kurtosis. For each case, Y is scaled so that its mean is
one. The raw-scale standard deviations varied from 1.14 (inverse Gaussian with log link) to 5.0 (lognormal with quadratic variance). The latter also revealed the largest skewness (19.06) and kurtosis (555)
on the raw-scale. As a preliminary assessment of the PV and QV model estimators, we examined the
mean and the standard deviation of λ̂, θ̂1 and θ̂2 across the 500 replicate data sets. The absolute bias for
the variance parameters was less than 0.011 while absolute bias for the link parameter was less than 0.09
√
in all cases. As expected, the standard errors estimated from samples of n = 10 000 were about 1/ 5
times the standard errors from samples estimated with n = 2000.
Figure 2 displays the % bias in µ̂(x), Dx (µ̂; x) and variance estimation for different estimators across
the midpoints of the deciles of X for selected DGPs (N = 2000) with link functions different from log.
EEE estimators correct the considerable bias in µ̂(x) evident in the GLM estimators with incorrect link
functions and appear to be consistent. The bias in Dx (µ̂; x) for GLM estimators increases at values of X
away from x = 0.5, whereas the EEE estimators are generally consistent across all values of X . Finally,
for gamma data with square-root link, even the gamma estimator with log-link shows biases in estimating
the variance. This arises due to the bias in estimating the mean function with incorrect link, since the
variance model is a function of the mean. Results for the three log-link processes and other non-loglink processes (not shown; results posted at http://home.uchicago.edu/~abasu/EEE/EEEWeb.pdf)
indicated that the EEE estimators were consistent for those cases. Results in Figure 2 were also similar
ˆ
for N = 10 000; the relatively small biases in µ̂(x), Dx (µ̂; x) and Var(Y
|X = x) with the EEE estimators
were further diminished.
Estimating marginal and incremental effects on health outcomes
Data: Gamma (sqr. root link), N = 2000
30
120
25
100
20
15
10
5
0
-5
Data: Gamma (sqr. root link), N = 2000
300
250
% Bias: var-hat
140
% Bias: dmudx-hat
% Bias: mu-hat
Data: Gamma (sqr. root link), N = 2000
35
101
80
60
40
20
0
200
150
100
50
0
- 20
-50
-10
- 40
-15
0.05 0.15 0.2 5 0 .35 0.4 5 0.5
x
0.05
0.65 0.75 0.85 0.95
0. 15 0 .2 5
x
Data: Inverse Gaussian (identity link), N = 2000
Data: Inverse Gaussian (identity link), N = 2000
10
4
2
0
-2
0.55
0.65
0.75
0.85
0 .95
150
% Bias: var-hat
% Bias: dmudx-hat
6
x
200
60
8
0.3 5 0 .45
Data: Inverse Gaussian (identity link), N = 2000
80
12
% Bias: mu-hat
- 100
- 60
0.05 0 .15 0.2 5 0 .35 0 .45 0.55 0.65 0. 75 0.85 0.95
40
20
0
-20
100
50
0
-50
-40
-4
-60
-6
0.05
0. 15
0.25
0 .35
0.4 5
x
0.5 5 0 .65
0.7 5
0.85 0 .95
Gamma
-100
0.05 0.15 0.25 0.3 5 0.45 0. 55 0.65 0.75 0.85 0.95
x
Poisson
Inv.Gauss
EEE.PV
0.05 0. 15 0.25 0.35 0.45 0 .5 5 0.65 0.75 0.85 0.9 5
x
EEE.QV
Fig. 2. % Bias in estimating µ, Dx (µ; x) and Var(Y |X = x) at different values of x from different estimators for
selected gamma and Inverse Gaussian data with different link functions.
Table 2 reports the bias and coefficient of variations in Dx (µ̂; x) at x = 0.2, 0.5, 0.8, and in ξ̂
arising for each combination of DGP and estimator. For gamma and inverse Gaussian data with log
link, the corresponding GLM estimators are maximum likelihood and hence are consistent and efficient
for Dx (µ; x) and for ξ . All three of the GLM estimators were consistent for the log-normal data with
v(x) = (1 + x). For these three log link processes, the EEE PV and QV estimators are consistent for
Dx (µ; x), though at a reduced efficiency that is the cost of estimating the link parameter λ. However,
the discrepancies in efficiency between GLM and EEE are considerably lower for estimation of ξ , which
averages Dx (µ; x) over X , and for the larger sample size.
Turning to three data generating mechanisms with non-log link (gamma with square root and inverse
link, and inverse Gaussian with identity link; Table 2), we see that misspecification of the link function
generally results in biased estimators of Dx (µ; x), especially at extreme values of x, and of ξ . The EEE
PV or QV estimators largely correct for these biases, and for ξ̂ do so with a high degree of efficiency.
For example, when a gamma with log link estimator is applied to gamma data with square root link
(N = 10 000), bias in estimating Dx (µ; x) is 28, 4 and 43% at x = 0.2, 0.5 and 0.8 respectively, and 17%
for ξ . Corresponding biases for EEE with either PV or QV structure are less than 0.5%, and the coefficient
of variation of ξ̂ is lower than for the GLM estimators.
For heteroscedastic log-normal data with quadratic variance, the true functional form of log{µ(x)} is
quadratic in x. However, as we only use a linear specification in x in our estimation models, all estimators
are expected to be biased for Dx (µ; x) and ξ . Interestingly, for this DGP, the EEE PV estimator overcomes
this problem by estimating a link parameter and produces an estimator with reduced bias over the GLM
estimators, although considerable sample size is required. We note that the estimating equations with QV
variance structure fail to converge for a significant number of replicates (80%) of the heteroscedastic
log-normal data. We believe that the variance structure imposed by QV may be very inappropriate for
102
A. BASU AND P. J. R ATHOUZ
Table 2. Simulation results on estimation of dµ/dx for alternative estimators on data generated with log
link (over all 500 replicates of N = 2000 and N = 10 000 each)
Data
Gamma
(log)
Inv.
Gaussian
(log)
% BIAS with datasets of N
Dx (µ̂; x)
Dx (µ̂; x)
at x = 0.2 at x = 0.5
Mn (cv)++
Mn (cv)
Gamma
−0.3 (8.3) 0.0 (11.3)
Poisson
−0.1 (8.6) 0.3 (11.8)
Inv. Gauss. 0.3 (9.1)
1.7 (13.8)
EEE (PV) 0.7 (26.7) −3.8 (12.8)
EEE (QV) 0.8 (26.6) −3.9 (12.8)
True
0.719
0.970
Estimator∗
Gamma
−0.1 (6.6)
Poisson
−0.1 (7.1)
Inv. Gauss. 0.0 (6.4)
EEE (PV) 1.5 (17.8)
EEE (QV) 2.9 (17.2)
True
0.705
Log
Gamma
normal
Poisson
V = (1 + x) Inv. Gauss.
EEE (PV)
EEE (QV)
True
= 2000
Dx (µ̂; x)
at x = 0.8
Mn (cv)
0.4 (14.6)
0.9 (15.1)
3.3 (19.1)
4.1 (29.0)
3.8 (28.7)
1.310
0.1 (8.9)
0.2 (11.2)
0.1 (9.6)
0.3 (12.2)
0.1 (8.7)
0.3 (11.0)
−1.8 (9.3) 1.5 (23.5)
−2.1 (9.1) −1.6 (21.9)
0.951
1.284
E x {Dx (µ̂; x)}
Mn (cv)
−0.1 (12.2)
0.3 (12.7)
2.0 (15.3)
4.8 (18.1)
4.6 (17.8)
1.013
% BIAS with datasets of N = 10 000
Dx (µ̂; x) Dx (µ̂; x) Dx (µ̂; x)
at x = 0.2 at x = 0.5 at x = 0.8
Mn (cv)
Mn (cv)
Mn (cv)
0.0 (3.6)
0.1 (4.9)
0.2 (6.3)
0.0 (3.8)
0.1 (5.2)
0.2 (6.6)
2.0 (6.1)
5.1 (11.7) 8.5 (18.1)
0.3 (11.5) −0.7 (5.2) 1.0 (13.6)
0.3 (11.5) −0.7 (5.2) 1.0 (13.6)
0.719
0.970
1.310
−0.1 (9.5)
0.0 (10.3)
0.0 (9.3)
2.0 (12.9)
0.4 (11.9)
0.993
0.0 (2.9)
0.0 (3.2)
0.0 (2.8)
0.2 (7.8)
0.8 (7.6)
0.705
0.0 (4.0)
0.1 (5.0)
0.0 (4.3)
0.1 (5.5)
0.0 (3.8)
0.0 (4.8)
−0.4 (3.9) 0.4 (10.7)
−0.5 (3.8) −0.6 (10.3)
0.951
1.284
0.1 (5.4)
0.1 (6.1)
0.1 (5.1)
−0.5 (5.2)
—
1.357
−0.3 (7.9) 0.0 (12.1)
−0.2 (8.5) 0.5 (13.8)
−0.5 (7.6) −0.3 (11.5)
−0.6 (21.4) −3.6 (12.2)
—
—
0.865
1.357
0.5 (16.8)
1.6 (19.8)
0.0 (15.9)
4.8 (31.7)
—
2.129
0.0 (14)
0.8 (16.2)
−0.5 (13.3)
5.7 (24.7)
—
1.493
0.0 (3.6)
0.0 (3.9)
0.0 (3.4)
0.3 (9.0)
—
0.865
42.4 (12.4)
27.7 (12.4)
71.0 (22.8)
2.3 (16.0)
1.0 (15.8)
2.520
16.9 (8.7)
7.1 (8.7)
36.7 (15.6)
1.1 (8.4)
0.5 (8.3)
1.924
−27.7 (1.2)
−28.8 (1.2)
−24.0 (1.4)
−0.1 (4.8)
0.1 (4.7)
1.320
Gamma
(sqr
root)
Gamma
Poisson
Inv. Gauss.
EEE (PV)
EEE (QV)
True
−27.7 (2.5) −3.7 (5.6)
−28.8 (2.7) −9.5 (5.9)
−23.9 (3.2) 8.2 (9.1)
−0.8 (10.9) −0.6 (6.4)
−0.1 (10.7) −0.9 (6.3)
1.320
1.920
Gamma
(inverse)
Gamma
Poisson
Inv. Gauss.
EEE (PV)
EEE (QV)
True
−24.7 (11.2)
−21.0 (12.1)
−26.1 (13.6)
−1.1 (20.6)
−1.0 (20.7)
−1.759
Inv.
Gaussian
(identity)
Gamma
Poisson
Inv. Gauss.
EEE (PV)
EEE (QV)
True
−26.4 (4.6) 1.1 (8.6)
39.2 (15.1)
−28.2 (4.9) −2.3 (9.0) 32.9 (15.5)
−24.2 (4.6) 5.6 (8.8)
47.2 (15.8)
−0.5 (12.0) −0.5 (10.1) 2.9 (21.7)
−0.3 (11.8) −0.3 (9.6) 2.8 (20.5)
1.000
1.000
1.000
8.8 (12.6) 32.8 (11.4)
12.7 (13.5) 36 (11.8)
6.7 (14.4)
30.3 (12)
−2.9 (15.4)
1.1 (25)
−2.9 (15.4)
1.1 (25)
−0.900
−0.545
−11.6 (11.0) −24.4 (4.8)
−8.1 (11.9) −21.1 (5.1)
−13.3 (12.9) −23.4 (11.6)
5.6 (22.4)
−0.2 (9.7)
5.7 (22.4)
−0.2 (9.7)
−1.156
−1.759
6.0 (9.7)
2.0 (10.1)
11.1 (10.1)
1.3 (9.0)
1.3 (8.7)
1.000
E x {Dx (µ̂; x)}
Mn (cv)
0.1 (5.3)
0.2 (5.6)
6.1 (13.6)
0.8 (6.6)
0.8 (6.6)
1.013
0.0 (4.2)
0.1 (4.6)
0.0 (4.1)
0.4 (5.4)
−0.1 (5.2)
0.993
0.2 (7.3)
0.3 (8.5)
0.2 (6.9)
0.9 (14.4)
—
2.129
0.2 (6.2)
0.2 (7.1)
0.2 (5.8)
0.9 (8.8)
—
1.493
−3.6 (2.5) 42.5 (5.5)
−9.7 (2.5) 27.0 (5.5)
8.0 (5.2) 70.3 (13.6)
−0.1 (2.9)
0.5 (7)
−0.1 (2.8) 0.1 (6.8)
1.920
2.520
17.2 (3.9)
6.9 (3.9)
36.5 (9.1)
0.2 (3.7)
0.1 (3.6)
1.924
9.5 (5.4)
13.1 (5.6)
9.5 (11.2)
−0.5 (7.4)
−0.5 (7.4)
−0.900
33.9 (4.8)
36.9 (4.9)
32.6 (7.6)
0.5 (11.6)
0.5 (11.6)
−0.545
−11.6 (4.7)
−8.4 (4.9)
−11.2 (10.3)
1.1 (8.1)
1.1 (8.1)
−1.156
−26.5 (1.9) 1.0 (3.6)
−28.2 (2.0) −2.6 (3.7)
−24.3 (2.0) 5.4 (3.8)
−0.2 (5.0) 0.0 (4.2)
−0.2 (5.0) −0.1 (4.2)
1.000
1.000
38.7 (6.4)
32.3 (6.5)
46.7 (6.8)
0.6 (8.4)
0.5 (8.5)
1.000
5.8 (4.1)
1.7 (4.2)
10.8 (4.3)
0.2 (3.7)
0.2 (3.7)
1.000
Log
Gamma
12.5 (10.6) 3.9 (18.4) −10.3 (24.8) −8 (22.3)
13.2 (5.2)
4.3 (8.8) −10.7 (11.4) −7.8 (10.3)
normal
Poisson
10.9 (9.4)
6.6 (21.8) −2.7 (38.1) −1.8 (33.3)
12.8 (4.8) 6.7 (10.8) −5.9 (16.8) −3.7 (14.8)
V = (1 + x)2 Inv. Gauss. 10.8 (9.9) −0.4 (14.8) −16.7 (18.2) −13.6 (16.6) 11.5 (4.7)
0.3 (6.9) −16.3 (8.2) −12.7 (7.6)
EEE (PV) 0.7 (19.8) −4.8 (14.6) 3.6 (41.2)
14.5 (64.1)
1.1 (8.1) −1.4 (6.7) 1.6 (19.8)
3.6 (20.4)
EEE (QV)
—
—
—
—
—
—
—
—
True
0.766
1.762
4.369
2.574
0.766
1.762
4.369
2.574
— We do not report the results obtained from this estimator on this data generating mechanisms due to convergence problems.
∗ Gamma, Poisson, and Inverse Gaussian estimators are implemented with log link.
++ Mn = Mean % Bias across 500 replicates (500−1 500
(Dx (µ̂; x)k − Dx (µ; x) · 100/Dx (µ; x)).
k=1
cv = coefficient of variation across 500 replicates (100 · sd(Dx (µ̂; x))/Dx (µ; x) ) where sd = standard deviation across 500
replicates.
True = the true value of the corresponding parameter.
Estimating marginal and incremental effects on health outcomes
103
modeling heteroscedastic log-normal data, resulting in instability in the weights for the mean model.
However, the EEE PV model converged in all replicates of the heteroscedastic log normal data and yielded
reasonably good fit in the mean and variance models.
4. E MPIRICAL EXAMPLE — HOSPITALIST STUDY
We now return to the hospitalist study described in the introduction to illustrate the proposed
methodology. Subjects are all adult patients (N = 6500) admitted to the medical wards at the University
of Chicago over a two-year period. Hospitalist and non-hospitalist attending teams rotated days through
the calendar in a fixed order. Thus, patients are assigned to attending physician in a quasi-random manner
based on date of admission, ensuring a balance of days of the week and months across the two sets of
attending physicians. Patients are clustered within physician and we will account for this clustering in our
variance estimation. There are no appreciable differences between the two groups of patients in terms of
demographics, diagnoses or other baseline characteristics. Inpatient (facility) expenditure is the outcome
of interest in our analysis, with a sample mean of $8530 (sd = $12500; 25th percentile = $2857; median
= $4910; 75th percentile = $9235). As discussed in the introduction, the two parameters of primary
interest are the incremental effect π1 of the hospitalist indicator variable (X 1 ) and the marginal effect of
disease-specific physician experience on total inpatient expenditure (Y ). Physician experience is measured
via the prior number of patients with the same disease treated by that physician (X 2 ). Other covariates that
are adjusted for include patient co-morbidities, relative utilization weight of diagnosis, admission month
indicator variables, and an indicator for transfer from another institution. Due to the skewed distribution
of the experience variable and to the non-linear relationship of expenditure to experience, and also to
conform to the specification used by the original investigators (Meltzer et al., 2002), we use logged count
of disease-specific experience as the covariate in our model. Since we are interested in the marginal effect
with respect to raw counts, we compute ξ2 as
∂µ(X )
∂µ(X )
1
ξ2 = E X
= EX
.
(4.1)
·
∂ X2
∂ log(X 2 ) X 2
We examine four models: (1) a gamma regression model with a log link, as was done in the original
study; (2) EEE model with PV variance structure; (3) a gamma regression model with square root link;
and (4) EEE model with the variance model fixed as Var(Y |X = x) = θ1 {µ(x)}0.5 . For each model,
we estimate the incremental effect π1 of the hospitalist variable as in (2.7), and the marginal effect ξ2 of
disease-specific experience via a natural adaptation of (2.6) to (4.1). To study the overall goodness of fit
of each of these models, we examine plots of the mean of the residuals (Yi − µ̂i ) across the fitted linear
predictor X iT β̂. Non-linear patterns in the residuals will reveal systematic bias in the fitted mean function
µ̂(x). We also present robust estimates of standard error based on (2.8) with A M computed as in (2.5) to
account for the fact that patients are clustered within attending physician. We also replaced the second
model with the EEE QV model with very similar results.
In Table 3, we present estimates of the incremental effect of the hospitalist variable and the marginal
effect of the disease-specific experience on the inpatient expenditure. Using gamma regression with log
link, the incremental effect of the hospitalist variable is estimated to be positive, while the EEE with
PV structure produces the sign and significance for this effect as expected from theory. Hospitalists
are expected to save costs; however, once physician experience is accounted for, cost savings are not
significantly different from zero. The marginal effect of experience, interpreted as the cost savings due
to the increase in disease-specific experience of one case averaged over all physicians, is evidently overestimated (indicating greater cost-savings with additional experience) with gamma log link estimator,
although the qualitative result is similar. The percentage bias in this case is estimated to be 22%
(= (318 − 260)/260).
104
A. BASU AND P. J. R ATHOUZ
Table 3. Estimated incremental and marginal effects on inpatient expenditures from hospitalist data
Mean $ (se+ )
Incremental effect of
Marginal effect of diseasespecific experience (ξ̂2 )
hospitalist (π̂1 )
Gamma (log link)
61 (279)
−318 (107)
EEE PV
−30 (231)
−261 (87)
Gamma (sqr root link)
−18 (215)
−242 (85)
122 (327)
−218 (128)
EEE with Var ∝ µ0.5
Note: models are adjusted for patient co-morbidities, relative utilization weight of
diagnosis, admission month indicator variables, and an indicator for transfer from
another institution.
+ Robust standard error based on sandwich estimator, accounting for clustering of
observations within physicians.
Estimator
The lack of fit for the gamma model with log link is illustrated in Figure 3(a), where the raw scale
residuals are plotted against the linear predictor. The gamma regression with log link (Figure 3(a)) tends
to over-predict at the right tail of the linear predictor, while the EEE with PV structure (Figure 3(b))
largely overcomes this problem in terms of predicted mean. Note, importantly, that the effects of even
modest poor fit of the mean function at the top deciles are amplified when the target is D2 (µ; x) or the
marginal or incremental effect. The effects of this lack of fit on estimation of D2 (µ; x) and subsequently
on ξ2 are illustrated in Figures 3 (c)–(e). Overestimation of µ in the right tail by the GLM log-link
model exaggerates D̂2 (µ; x) and thereby biases ξ̂2 . The EEE–PV provides a better fit to the data across
all values of x, and in addition yields a model in which D2 (µ; x) is less strongly related to the linear
predictor, leading to less biased estimation of ξ2 . Figure 3(d) shows how the EEE–PV model corrects the
overestimated D̂2 (µ; x) in the GLM log-link model.
The estimated link parameter under the PV variance structure is λ̂ = 0.398 (95% CI: 0.191, 0.605),
indicating that the optimal link for these data is not the log link but more likely a square root link. The PV
model fit also suggests that the variance is a quadratic function of the mean (θ̂1 = 0.81, 95% CI: 0.71, 0.91
and θ̂2 = 2.14, 95% CI: 1.97, 2.31), which suggests that (Y |X ) is close to a gamma distribution. Using
this information, we ran the standard gamma GLM model, this time with a square root link. This estimator
(Table 3) produces marginal effects that are more in line with the EEE model. Note the mild increase in
efficiency due to using a known link function.
Finally, to see the efficiency benefit of flexibly modeling the variance using EEE, we compared our
results to a modified EEE estimator where we incorrectly fixed the variance to be proportional to {µ(x)}0.5 ,
but allowed the link parameter λ to be estimated. Though this estimator is consistent, the standard errors
for the incremental and the marginal effects are much larger than when a flexible variance model is
used (Table 3). Had we used this estimator, we would have concluded that the marginal effect of disease
experience is not statistically significant at the 5% level.
5. D ISCUSSION
In this paper, we have proposed estimating equations for parameters in the link and variance functions
along with those of the linear predictor in a generalized linear model, and have developed methodology
for using this fitted model to estimate additive marginal and incremental effect parameters. Our method
addresses difficulties in choosing the correct link and variance functions in these models. The work is
important since, in many health applications, researchers are primarily interested in estimating the mean
and functionals of the mean of the outcome variable in the original versus a transformed scale, and effect
50000
–1
0
1
Linear Predictor
2
3
–1
0
2
GAMMA - Log link Model
(c)
3
4
0
–6000
GLM-Log
EEE PV
Estimators
(d)
–8000
0
1
Linear Predictor
3
dmu/d(Experience)
–4000
–2000
0
Predicted dmu/d(Experience)
–6000
–4000
–2000
–1
1
2
Linear Predictor
EEE-PV model
(b)
–8000
–8000
–6000
dmu/d(Experience)
–4000
–2000
0
GAMMA - Log link Model
(a)
–2
105
Residuals/Lowess smoother
0
–2
–50000
–50000
Residuals/Lowess Smoother
0
50000
Estimating marginal and incremental effects on health outcomes
–1
0
1
2
Linear Predictor
3
4
EEE-PV model
(e)
Fig. 3. (a) Scatter plot of raw-scale residuals against the linear predictor from the GLM log-link model. (b) Scatter
plot of raw-scale residuals against the linear predictor from the EEE PV model. (c) Plot of D̂Experience (µ; x) versus
the linear predictor from the GLM log-link model. (d) Comparison of D̂Experience (µ; x) from the GLM log-link
model to D̂Experience (µ; x) from the EEE–PV model. (e) Plot of D̂Experience (µ; x) versus the linear predictor from
the EEE–PV model.
estimates obtained are additive even though the scale of estimation of these effects is data dependent,
being determined through the link function.
We develop our methodology using two commonly used functionals of the mean—the marginal and
the incremental effects. However, our methodology would apply equally well to the variants of these effect
parameters. For example, a related but conceptually different definition of marginal effect for continuous
covariates is E X − j {D j (µ; X )| X j =x j }. Here D j (µ; X ) is evaluated at a specific value of X j = x j , and
then the expected value is calculated over the distribution of X − j , marginally with respect to X j . An
example is the effect of price (X j ) on the expected demand of a good (Y ) averaged over all characteristics
of consumers (X − j ) and evaluated at the current price (x j ). Other functionals of µ(x) of interest include
E X {D j (µ; X )/µ(X )} and E X {D j (µ; X ) · X j /µ(X )}. The former represents the average proportionate
change in µ(X ) corresponding to a unit change in X j in the whole population, adjusted to the population
distribution of X . The latter is the well-known elasticity measure in economics representing the population
average percentage change in the mean of Y associated with a one percent change in X j . When
X j is
binary, another measure of incremental effect (Oaxaca, 1973) is given by E X − j |X j =1 D j (µ; X ) , where
the expected value is over X − j , conditional on X j = 1. It represents the effect of X j on µ(X ), adjusted to
the distribution of X − j in the population with X j = 1. Examples include the effect of prevention of heart
attacks in the population of patient who have heart attacks and providing insurance to people who do not
have health insurance. The choice of which if any of these adjusted effect measures is of interest depends
on the specific research question being addressed.
106
A. BASU AND P. J. R ATHOUZ
Other researchers have proposed methods for estimating link and/or variance functions along with
regression coefficients in GLMs. Nelder and Pregibon (1987) suggest a profile extended quasi-likelihood
function that iteratively estimates regression parameters β and ancillary parameters in the link and
variance functions. Scallan et al. (1984), Mallick and Gelfand (1994) and Kaiser (1998) propose maximum
likelihood estimation of parametric link function models, but their approaches require full distributional
assumptions, and neither Scallan et al. nor Kaiser include any simulation studies to assess the performance
of their estimators. Other approaches include quasi-likelihood methods where a variance function is
estimated but the link function is assumed known (Chiou and Müller, 1998), where the link function is
estimated non-parametrically but the variance function is assumed known (Li and Duan, 1989; Li, 1991;
Weisberg and Welsh, 1994; Carroll et al., 1997), and where non-parametric link and variance functions
are estimated simultaneously (Chiou and Müller, 1998). None of these researchers have applied these
methods to the estimation of marginal or incremental effects.
A critique often leveled against methods that involve estimation of a link function is that, as the
link function varies, so does the interpretation of regression coefficients β j . This is indeed a problem
when the primary focus is on β j . However, the parameters of interest here are the mean function µ(x)
and the marginal and incremental effects whose interpretation is invariant to link function. The flexible
link function allows for less biased estimation of µ(x) across a range of values of x, thereby improving
estimation of effect measures that are functionals of µ(x). Moreover, the advantage of this approach lies
to a considerable extent in overcoming the problems of lack of fit without changing the linear specification
of covariates.
Evidence from simulations shows that these estimators, especially those with PV variance structure,
perform well in terms of bias and efficiency when the distribution of the outcome variable is not known
and/or there is ambiguity about the appropriate link function. Estimation of the link parameter λ does
incur a cost in terms of efficiency, but this is partially recovered through simultaneous estimation of the
variance structure. One surprising result was that, while relative efficiency losses due to link function
estimation were sometimes substantial for effects Dx j (µ; x) for given x, corresponding losses were much
more modest for integrated effects ξ j . In applications, we recommend the use of PV variance structure
for continuous outcome variables such as costs and expenditures since the gamma and inverse Gaussian
distributions are special cases of this variance structure. Similarly, we recommend the use of QV variance
structure for discrete outcomes such as length of stay and counts of physician visits since Poisson and
negative binomial distributions are its special cases. Practically, we also found that the new estimators
work best in analyses with larger sample sizes, say over N = 5000, sample sizes that are not uncommon
in health economics and health policy applications.
Finally, we mention generalized additive models (Hastie and Tibsharini, 1990) as an alternative to
the model we propose with flexible link function. While useful in some contexts, in many applications,
there are several covariates, and fitting such models with multiple smooth terms becomes difficult. The
flexible yet parametric link function approach that we propose offers an added degree of flexibility
over the standard generalized linear model, while retaining enough model structure so that the model
is still relatively easy to estimate. We hope that this methodology will be increasingly used in the health
economics and other areas of research that are plagued by data characteristics that makes a priori choices
of link functions and of estimators with distributional assumptions difficult.
ACKNOWLEDGEMENTS
We are grateful to Willard G. Manning, Ronald A. Thisted, Vanja Dukic, and Daniel L. Gillen for
extremely helpful comments and suggestions. We would also like to thank the editors and referees for
their valuable comments that have improved the manuscript substantially. The authors also thank David
Estimating marginal and incremental effects on health outcomes
107
Meltzer for providing the Hospitalist Study data. The opinions expressed are those of the authors, and not
those of the University of Chicago. Anirban Basu’s time was supported in part by the National Institute
on Alcohol Abuse and Alcoholism (NIAAA) grant 1 RO1 AA12664-01 A2.
APPENDIX A
Sketch of proof
The proof follows standard asymptotic arguments for solutions to estimating equations. Let γ0 be
the true value of γ = (β1 , β2 , . . . , β p , λ, θ1 , θ2 )T and let γ̂ = (β̂√1 , β̂2 , . . . , β̂ p , λ̂, θ̂1 , θ̂2 )T solve
G γ = 0. Then, under regularity conditions, one can approximate N (γ̂ − γ0 ) via Taylor series
−1 N
N
approximation by −N −1 i=1
∂G iγ /∂γ
G iγ + o p (1). By the multivariate central limit
N −0.5 i=1
N
N
L
theorem, N −0.5 i=1
G iγ −→M V N (0, B∞ ), where B∞ = lim N →∞ N −1 i=1
E(G iγ G iγT ). By the
p
N
N
law of large numbers, N −1 i=1
∂G iγ /∂γ −→ lim N →∞ N −1 E( i=1
∂G iγ /∂γ ) = C∞ . Hence, C∞ =
N
N
N −1 E( i=1
∂G iγ /∂γ ) + o(1) and N −1 i=1
∂G iγ /∂γ = C∞ + o p (1). Therefore by Slutsky’s theorem,
√
−T
N (γ̂ − γ0 ) has the limiting distribution of M V N (0, A∞ ) where A∞ = C−1
∞ B∞ C∞ . A N in (2.4),
computed replacing γ with γ̂ , yields a consistent estimator of A∞ .
APPENDIX B
Initial values of the regression coefficients come from the estimates of regression coefficients from
a gamma GLM model with log link. The initial value of the link parameter λ is set to 0.1. For the PV
structure, initial value of θ1 comes from the shape parameter (ϕ) computed by the gamma GLM model.
The initial value of θ2 comes from the modified Park test (Manning and Mullahy, 2001). In this test the
logarithm of the squared residuals from the GLM model is regressed on the logarithm of the predicated
values (µ̂) from the GLM model. The coefficient of the log(µ̂) gives initial estimate for θ2 . For the QV
structure, the squared residuals from the GLM model is regressed on the predicated values (µ̂) and the
squared predicted values (µ̂2 ) without an intercept. The coefficient of µ̂ gives an estimate for θ1 and the
coefficient of µ̂2 gives an estimate for θ2 .
(k)
Parameter estimates are updated using the following equality: γ̂ (k+1) = γ̂ (k) + I(k)−1 G γ , where
(k)
(k)
I (k) = E(−∂G γ /∂γ ). I (k) and G γ are computed using the current value of γ̂ (k) . This procedure is
iterated until the maximum relative difference in parameter estimates between two successive iterations
γ (k) and γ (k+1) is less than 0.0001.
We ensure that the required condition µi > 0 ∀i is met by setting µ̂i to missing for all observations
for which this condition is violated at any given iteration. After the estimator has converged we searched
for any observation with missing µ̂i . We did not find any such observations for all our simulated datasets
and also for the empirical example with the hospitalist data.
APPENDIX C
N
i , X = 1) − µ(X i , X = 0)} be an estimator
Let π̃ j = Ê X − j {D j (µ; X − j )} = N −1 i=1
{µ(X −
ij
ij
j
−j
for π j when γ is a vector of known constants. When γ is estimated from the data, then the estimator
N
i , X = 1) − µ̂(X i , X
for π j is π̂ j = N −1 i=1
{µ̂(X −
ij
i j = 0)}. Following a first-order Taylor
j
−j
108
A. BASU AND P. J. R ATHOUZ
series approximation, one can write π̂ j = π̃ j +
N
N
N −1 i=1
∂G iγ /∂γ and Sγ = N −1 i=1
G iγ . Thus,
Var π̂ j
∂π j ∂γ γ (γ̂
− γ ) where (γ̂ − γ ) = −1
γ Sγ , γ =
T
N ∂π j ∂π j ∂π j −1
−1
i iT
−T
(γ )
( )(π̃ j − π̄ j )Sγ +
Sγ Sγ (γ )
= Var π̃ j + 2
∂γ γ γ
∂γ γ
∂γ γ
i=1
T
∂π j ∂π j ∂π j −1
AN
= Var π̃ j + 2
( )(π̃ j − π̄ j )Sγ +
∂γ γ γ
∂γ γ
∂γ γ
N
N
where Var π̃ j = N −1 (N − 1)−1 i=1
(π̃ j −π̄ j )2 , π̄ j = N −1 i=1
π̃ ji and A N is the analytical
variance of parameter as in (2.4). Since E γ {Sγ } = 0, the second term in the above expression converges
to zero. Note that the second part is identical to the delta method used to calculate the asymptotic variance
of a non-linear function of parameters.
The variance of the marginal effect can also be obtained through an analogous method.
R EFERENCES
BASU , A., M ANNING , W. G. AND M ULLAHY , J. (2002). Comparing alternative models: log vs proportional hazard?
Forthcoming in Health Economics.
B LOUGH , D. K., M ADDEN , C. W. AND H ORNBROOK , M. C. (1999). Modeling risk using generalized linear models.
Journal of Health Economics 18, 153–171.
B OX , G. E. P. AND C OX , D. R. (1964). An analysis of transformations. Journal of the Royal Statistical Society B
26, 211–252.
B RADLEY , E. AND T IBSHIRANI , R. (1993). An Introduction to the Bootstrap. New York: Chapman and Hall.
C ARROLL , R. J.
Hall.
AND
RUPPERT , D. (1988). Transformation and Weighting in Regression. New York: Chapman and
C ARROLL , R. J., FAN , J., G IJBELS , I. AND WAND , M. P. (1997). Generalized partially linear single-index models.
Journal of the American Statistical Association 92, 477–489.
C HIOU , J. M. AND M ÜLLER , H. G. (1998). Quasi-likelihood regression with unknown link and variance functions.
Journal of the American Statistical Association 98, 1376–1386.
C HIOU , J. M. AND M ÜLLER , H. G. (1999). Nonparametric quasi-likelihood. Annals of Statistics 27, 36–64.
C ROWDER , M. (1987). On linear and quadratic estimating functions. Biometrika 74, 591–597.
D UAN , N. (1983). Smearing estimate: a nonparametric retransformation method. Journal of the American Statistical
Association 78, 605–610.
G REENE , W. H. (2000). Econometric Analysis, 4th edn. Englewood Cliffs, NJ: Prentice-Hall.
H ALL , D. B. AND S EVRINI , T. A. (1998). Extended generalized estimating equations for clustered data. Journal of
the American Statistical Association 93, 1365–1375.
H ASTIE , T. J.
AND
T IBSHARINI , R. J. (1990). Generalized Additive Models. London: Chapman and Hall.
H OSMER , D. W. AND L EMESHOW , S. (1995). Applied Logistic Regression, 2nd edn. New York: Wiley.
H UBER , P. J. (1972). Robust statistics: a review. The Annals of Mathematical Statistics 43, 1041–1067.
K AISER , M. S. (1998). Maximum likelihood estimation of link function parameters. Computational Statistics and
Data Analysis 24, 79–87.
Estimating marginal and incremental effects on health outcomes
109
L I , K. (1991). Sliced inverse regression for dimension reduction. Journal of the American Statistical Association 86,
316–327.
L I , K. AND D UAN , N. (1989). Regression analysis under link violation. Annals of Statistics 17, 1009–1052.
L IANG , K. Y.
73, 13–22.
AND
Z EGER , S. L. (1986a). Longitudinal data analysis using generalized linear models. Biometrika
L IANG , K.-Y.
73, 13–22.
AND
Z EGER , S. L. (1986b). Longitudinal data analysis using generalized linear models. Biometrika
M ALLICK , B. K. AND G ELFAND , A. E. (1994). Generalized linear models with unknown link functions. Biometrika
81, 237–245.
M ANNING , W. G. (1998). The logged dependent variable, heteroscedasticity, and the retransformation problem.
Journal of Health Economics 17, 283–295.
M ANNING , W. G. AND M ULLAHY , J. (2001). Estimating log models: to transform or not to transform? Journal of
Health Economics 20, 461–494.
M ANNING , W. G., BASU , A. AND M ULLAHY , J. (2003). Generalized modeling approaches on risk adjustment to
skewed outcomes. NBER Working paper. No. t0293.
M CCULLAGH , P. (1983). Quasi-likelihood functions. Annals of Statistics 11, 59–67.
M CCULLAGH , P. AND N ELDER , J. A. (1989). Generalized Linear Models, 2nd edn. London: Chapman and Hall.
M ELTZER , D. O., M ANNING , W. G., M ORRISON , J., G UTH , T., H ERNANADEZ , A., D HAR , A., J IN , L. AND
L EVINSON , W. (2002). Effects of physician experience on an academic general medicine service: results of a trial
of hospitalists. Annals of Internal Medicine 137, 866–874.
M ULLAHY , J. (1998). Much ado about two: reconsidering retransformation and the two-part model in health
econometrics. Journal of Health Economics 17, 247–281.
N ELDER , J. A. AND P REGIBON , D. (1987). An extended quasi-likelihood function. Biometrika 74, 221–232.
OAXACA , R. L. (1973). Male-female wage differentials in urban labor markets. International Economic Review 14,
693–709.
P REGIBON , D. (1980). Goodness of link tests for generalized linear models. Applied Statisics 29, 15–24.
P REGIBON , D. (1981). Logistic regression diagnostics. Annals of Statistics 9, 705–724.
S CALLAN , A., G ILCHRIST , R. AND G REEN , M. (1984). Fitting parametric link functions in generalised linear
models. Computational Statistics and Data Analysis 2, 37–49.
S TATACORP., (2001). Stata Statistical Software: Release 7.0. College Station, TX. Vol. 2, p. 406.
W EDDERBURN , R. W. M. (1974). Quasi-likelihood functions, generalized linear models, and the Gauss–Newton
method. Biometrika 61, 439–447.
W EISBERG , S.
AND
W ELSH , A. H. (1994). Adapting for the missing link. Annals of Statistics 22, 1674–1700.
[Received April 3, 2003; first revision September 22, 2003; second revision June 28, 2004;
accepted for publication June 30, 2004]