TECHNOLOGICAL IMPROVEMENT AND THE - Berkeley-Haas

TECHNOLOGICAL IMPROVEMENT AND THE DECENTRALIZATION PENALTY IN A
SIMPLE PRINCIPAL/AGENT MODEL
Ruochen Liang∗ and Thomas Marschak∗∗
April 10, 2017
We consider the organizers of a firm who compare a decentralized arrangement where
divisions are granted total autonomy with an arrangement where perfect monitoring and
policing guarantee that all divisions make the choices that management wants them to
make. We ask: when does improvement in the divisions’ technology strengthen the case
for decentralization and when does it weaken it? The question is difficult and it is natural
to start with a stripped-down model. We do so by considering a Principal and a single
Agent. When x is the Agent’s effort, the organization achieves the surplus R(x)−t·C(x),
where R is revenue, t is a positive technology parameter known to both parties, and t·C(x)
is the cost of the Agent’s effort. When technology improves, t drops. In the decentralized
mode the Agent autonomously chooses x, bears the resulting cost t · C(x),and receives a
share of the resulting revenue, namely r · R(x), were 0 < r < 1. The Principal receives
the residual (1 − r) · R(x). In the exogenous case the share r is specified outside the
model (it might be the result of Principal/Agent bargaining). In the endogenous case
the Principal, who knows how the Agent responds to every possible r given the current
t, chooses the r which maximizes the residual revenue. The Decentralization Penalty for
a given t equals the maximal possible surplus for that t minus the surplus achieved in the
decentralized mode. It turns out that there are no simple conditions on C and R which
imply that the Penalty grows (shrinks) when technology improves. Instead, we obtain a
variety of results about relations between t, surplus, the Principal’s “generosity” (the size
of the principal’s chosen share in the endogenous case), and “effectiveness” (the effect of
a small rise in the share on the Agent’s effort).
* Department of Mathematics, University of California, Berkeley
** Walter A. Haas School of Business, University of California, Berkeley.
1
TECHNOLOGICAL IMPROVEMENT AND THE DECENTRALIZATION PENALTY IN A
SIMPLE PRINCIPAL/AGENT MODEL
1. Introduction
Does the case for decentralizing a firm get stronger or weaker when the technology used by
one or more of its divisions improves? Management seeks a good balance between the cost of the
divisions’ efforts and the revenue which those efforts yield. One way to achieve the balance may
be intrusive and expensive monitoring and policing, ensuring that each division’s effort is what
Management wants it to be. A better way might be “decentralization”: grant autonomy to each
division, let it bear the costs of its chosen effort, and reward it when revenue is realized, using
a reward formula which the division accepts and which Management prefers among all formulae
that the division accepts. How does a drop in the cost of effort, due to improved technology,
affect the relative performance of the two modes of organizing?
The question could be posed in a full Principal/Agents model of the decentralized mode.
The Agents’ effort choices might be hidden. There could be a random state variable which
affects the revenue that a given effort yields, or the cost of a given effort, or both. Technological
improvement may change the distribution of the state variable. The Principal might know less
about the current state than the Agents, might seek information about what they know, and
wants a truthful response. If there is more than one Agent, each Agent’s earnings may depend on
the effort choices of all. Then the question to be studied is the effect of technological improvement
on the equilibria of the game.
In the abundant firm-design and Principal/Agent literature the model often includes a cost
of effort. But the effect of a drop in that cost on the merits of decentralization is seldom the
central question. As we just saw, the question is complicated in a full Principal/Agents model.
Yet the question has strong motivation. The ongoing IT revolution, for example, calls for better
understanding of one important case: the Agent is an information-gatherer who sends a signal
about the current state to the Principal. Revenue depends on a choice that the Principal makes
when the signal is received, and more effort by the Agent yields a more useful signal. Improvement
in IT lowers the cost of any given effort. Does that strengthen or weaken the case for Agent
autonomy?1
2. The model
We shall consider a drastically stripped-down Principal/Agent model in which the effect
of technology improvement on the merits of decentralization can nevertheless be studied. We
consider a Principal and a single Agent. There is no informational asymmetry. The Agent is
1
In particular, the Principal might choose a production quantity in response to the Agent’s signals. That
situation is studied in Marschak, Shanthikumar, and Zhou (2017). The question that paper asks is whether the
average production quantity rises or falls when the Agent’s information-gathering technology improves (or the
Agent works harder). If improved technology indeed causes average production quantity to rise, and if a rise in
quantity implies a rise in surplus, then Agent autonomy becomes more attractive when technology improves.
2
autonomous and freely chooses his effort, but it is not essential to the model that his choice be
hidden. The Agent agrees to a fixed share of realized revenue as his reward. The model is so
simple that one might expect unambiguous answers to our question to emerge very quickly. We
shall see that this is not the case and that even our stripped-down model yields a surprisingly
rich array of propositions about the effect of technological improvement on the performance of
the decentralized mode. Some of them are not suggested by any simple intuition.
Formally, the Agent’s effort x belongs to a set Σ ⊆ IR+ of possible positive efforts. The
effort x generates a positive revenue R(x), where R is strictly increasing. The effort x costs the
Agent t · C(x), where C is positive and strictly increasing. A drop in t occurs when the Agent’s
technology improves (or when there is a fall in the price of the inputs which the Agent’s effort
requires). The functions R and C, and the parameter t, are known to both parties. We shall
consider the surplus at the effort x, denote W̃ (x, t). Thus
W̃ (x, t) = R(x) − t · C(x).
Surplus may also be called profit. It equals the net amount earned by the Principal (the residual
revenue) plus the net amount earned by the Agent.
High surplus is desired by the firm and by “society” as well. In an alternative setting there is
an Owner, who rewards the Principal with a share — say α ∈ (0, 1) — of profit, and allows the
Principal to choose between the two modes of organizing. Then the Principal wants to achieve
a high value of α · W̃ (x, t). Introducing the constant α does not change our main conclusions.
In the centralized mode, the Principal monitors and polices the Agent perfectly to insure that
the Agent exerts the effort which maximizes surplus (profit). In the decentralized mode there is
no direct monitoring. The Agent freely chooses x ∈ Σ, bears the cost t · C(x), and is rewarded for
his effort by the Principal. We study an extremely simple reward scheme, namely linear revenue
sharing. The Principal observes the current revenue and rewards the agent with a share r ∈ [0, 1]
of the revenue. If the Agent chooses the effort x, then the Agent earns rR(x) − t · x and the net
amount received by the Principal is (1 − r) · R(x).2 The Agent chooses to exert the effort x̂(r, t),
the smallest maximizer of rR(x) − t · x on the set Σ.3 We denote the surplus when the share is
r by W (r, t) (the tilde is deleted). So
W (r, t) ≡ W̃ (x̂(r, t), t) = R(x̂(r, t)) − t · C(x̂(r, t)).
Note that if r = 1, then the Agent’s effort choice x̂(1, t) is surplus-maximizing. Thus
W (1, t) = W̃ (x̂(1, t), t) is the largest possible surplus.
In the centralized mode, monitoring and policing insure that W (1, t) is achieved.
2
We could add a participation constraint. To be willing to participate, the Agent requires a total payment of
at least Q. If the Principal pays Q regardless of effort, then at effort x the Agent receives Q + r · R(x) − t · C(x)
and the Principal receives (1 − r) · R(x) − Q. Incorporating the constant transfer Q into our analysis would not
change our results in an essential way.
3
Our results would be essentially unchanged if we studied the largest maximizer instead.
3
We shall study the decentralized mode by considering two cases. In the exogenous case, the
reward share is determined outside the model. It might, for example, be the result of previous
bargaining between Principal and Agent, or it might be prescribed by law. In the endogenous
case, the reward share is chosen by the Principal so as to maximize the residual (1−r)·R(x̂(r, t)),
the portion of revenue that the Principal keeps, given that the Agent uses the function x̂ when
he responds to a given share. We let r∗ (t) denote the share the Principal chooses. So in the
endogenous case the Agent’s effort is x̂(r∗ (t), t) and surplus is W̃ (x̂(r∗ (t), t), t) = W (r∗ (t), t).
Note that in the model and its two decentralized cases, all shares r in [0, 1] are possible. But
the effort set Σ, and the set of possible values of the technology parameter t, might be finite or
they might be continua. In studying particular examples, both finite and nonfinite, it is often
helpful to let the open interval (0, 1) be the set of shares considered by the Principal.
One of our main concerns is the Decentralization Penalty. We shall often use surplus gap as
an alternative term for Penalty. The Penalty, as we define it, is the difference between maximal
surplus — which perfect monitoring and policing achieves — and the surplus achieved in the
decentralized mode. An alternative definition of penalty would be the ratio of decentralized
surplus to maximal surplus. Both the difference and the ratio are worth studying. The difference
concept is natural if one knows the cost of perfect monitoring and policing. That cost drops
when monitoring technology improves. If improvement in the Agent’s technology substantially
raises (lowers) the surplus difference, then perfect monitoring and policing become more (less)
attractive. Studying the Penalty, as we define it, allows us to trace the organizational implications
of improvements in both kinds of technology. Our results about the Agent’s technology and the
Penalty (defined as a difference) turn out to be surprisingly complex and varied. It appears
technically difficult to obtain parallel results when the Penalty is a ratio, and there is no reason
to believe that such parallel results would be less complex and less varied than the ones we
obtain. In the Related Literature section below we comment on computer-science studies of the
ratio (the “Price of Anarchy”). In our Concluding Remarks section we suggest the ratio as one of
a number of ways of extending our results and we briefly point to some examples where looking
at the ratio leads to results that contrast sharply with ours.
Our Decentralization Penalty is W (1, t) − W (r, t) in the exogenous case and W (1, t) −
W (r∗ (t), t) in the endogenous case. We shall see, in many examples, that if t drops, then both
the first term W (1, t) and the second term W (r, t) (or W (r∗ (t), t)) rise. Then, if we want to find
the effect of a drop in t on the Penalty, we face a challenging question: which of the two terms
rises faster when t drops?
In addition to the effect of a drop in t on the Penalty (the surplus gap), we study the following
questions:
• In each case — exogenous and endogenous — what is the effect of technological improvement
— i.e., a drop in t — on the Agent’s effort? Does the Agent always work harder when
technology improves?
• In both cases, is technological improvement always “good news” for the Agent? For the
Principal? Is it good news from the welfare point of view?
4
• We study the effectiveness of a share increase — the increase in effort due to a small increase
in the share. Does effectiveness rise or fall when t drops? That is an important question
in both cases.
• In both cases we consider the effort gap — the amount by which effort under decentralization
falls short of surplus-maximizing effort. That gap is x̂(1, t) − x̂(r, t) in the exogenous case
and x̂(1, t) − x̂(r∗ (t), t) in the endogenous case. Does the gap rise or fall when technology
improves?
• In many examples the effect of technology improvement (a drop in t) on the effort gap is easy
to determine.(Typically we only need to find the sign of the cross partial x̂rt ). The effect
of technology improvement (a drop in t) on the surplus gap may be much more difficult.
Accordingly we ask another question, which is itself interesting: when does the surplus
gap track the effort gap? When does a drop in t cause the two gaps to move in the same
direction? The answer is useful to the Principal if he wants to check whether a technical
advance weakens or strengthens the case for remaining decentralized rather than moving
to perfect monitoring. It then suffices for the Principal (who knows the old and new values
of t) to observe the Agent’s effort for the new value of t and to compare it with the new
surplus-maximizing effort x̂(1, t).
• In the endogenous case, does the Principal become more or less generous when technology
improves? (Does the share r∗ (t) rise or fall when t drops?).
• The “tracking” question is difficult, especially in the endogenous case in many examples.
We study it by considering its relation to effectiveness and generosity. Each of these may
increase or decrease when there is a drop in t, so there are four possible combinations. We
shall examine the tracking question for each of those combinations.
3. Related literature.
To make a connection between our problem and standard Principal/Agent results, consider
the simplest moral-hazard framework, adapted to fit our problem. The Agent has two effort
choices, xL and xH , where 0 < xL < xH . The lower effort costs the Agent t · CL , while the
higher effort costs t · CH , where t > 0 is our technology parameter, known with certainty by both
parties. Let the outcome of an effort be a random variable. It is one of two revenues RL and
RH , where 0 < RL < RH . The effort xL yields RH with probability q and RL with probability
1 − q. The effort xH yields RH with probability p and RL with probability 1 − p, where p > q.
If the Agent declines to participate, he receives a reservation amount which we normalize to
be zero. The Principal proposes a contract, which the Agent accepts. Under the contract, the
Agent receives wL from the Principal if revenue turns out to be RL and wH if revenue turns
out to be RH , where wH > wL and wL may be negative. Suppose that both parties are riskneutral and there is no liability constraint on |wL |. Then among all contracts acceptable to
the Agent, the Principal’s favorite one induces the Agent to make the choice that maximizes
5
average surplus.4 Let our Decentralization Penalty be the difference between highest attainable
(“first best”) average surplus and average surplus under the Principal’s favorite contract. With
no liability constraint, the Penalty is zero, whatever t may be. What we have called the “effort
gap” is also zero (the Agent’s choice is “first best”).
When there is a liability constraint, the Agent might not have the assets to cover wL and
hence the Principal’s menu of contracts narrows. The Penalty and the effort gap may now be
positive. The relevant literature is large, but as far as we can determine, it lacks a systematic
study of the effect of a drop in the Agent’s costs on the Penalty and on the effort gap. That
may be because, as we have seen, such a study would be difficult. It becomes easier if we further
constrain the Principal’s available contracts. The Principal is no longer free to choose any
assignment of rewards to observed outcomes. He is constrained, in our version of the problem,
to use linear sharing.
It remains true, of course, that a great many papers, starting with the earliest ones, consider
first-best surplus maximization and first-best Agent effort as benchmarks which one can compare
with surplus and effort under the Principal’s chosen contract. In many of these papers, there
is a cost for every Agent effort, and it is the Agent who bears the cost.5 Under reasonable
conditions, a drop in cost increases both first-best and decentralized surplus and both first-best
effort and decentralized effort. Unfortunately that does not tell us the effect of a drop on the
surplus difference and on the effort difference.
If we allow more than one Agent, then parts of the large literature on the design of organizations become relevant. The designer has a goal, say surplus (profit) maximization, and can
choose between a structure where a single member commands the choices made by all the others,
and a structure where everyone is autonomous. The latter structure might be modeled as a game.
A rather small piece of the design literature6 studies the communication and computation costs
of each structure and the trade-off between those costs and some measure of gross performance
(e.g., gross expected surplus, before the costs are subtracted). The problem is far more complex
than the one we consider here and the results remain scarce and specialized. 7
Finally, we note a related strand of research by computer scientists, which they label “the
price of anarchy”. Typically the object of study is a game and the price of anarchy is the ratio of
the payoff sum in the socially “worst-case” equilibrium to the highest attainable sum. A simple
example is a symmetric two-player Prisoners’ Dilemma, where the price of anarchy is the sum of
4
See, for example, Chapter 4 in Laffont and Martimort (2002).
One of the earliest of these papers is Holmstrom (1979), where the Agent’s utility depends on his wealth w
and on his action a, and takes the form U (w) − a. If every action has a cost, borne by the Agent, and if there is
a one-to-one correspondence between actions and costs, then we can interpret a itself to be the Agent’s cost.
6
Surveys of this piece of the literature are found in Garicano and Prat (2011) and Marschak (2006).
7
In Courtney and Marschak (2009) a “sharing game” is studied. There are n autonomous players. Each
chooses an effort and bears its cost, and the n choices determine a revenue which all share. An equilibrium of
the game may be efficient (surplus is maximized), shirking (the chosen efforts are less than the efficient ones), or
squandering (the efforts are greater than the efficient ones). It is shown, under differentiability assumptions, that
if we are in a squandering equilibrium, if the players’ costs shift downward, and if the new equilibrium is again
a squandering one, then the Decentralization Penalty at the new equilibrium (the amount by which surplus falls
short of its maximum) is less than the Penalty at the old one. That need not be true for shirking equilibria.
5
6
payoffs in Nash equilibrium divided by the sum of payoffs when the players cooperate. A variety
of social situations are studied from this point of view.8 Many of these studies develop bounds
on the price of anarchy. It appears, however, that up to now this literature has not examined the
Principal/Agent setting. In our Conclusion we shall make some observations about what might
happen if we redefined our Decentralization Penalty so that it becomes a ratio rather than a
difference.
4. Plan of the remainder of the paper.
In Section 5 we examine four examples where the set of possible efforts and the set of possible
values of t are not finite. The examples provide a preview of our general results. In Section 6 we
develop basic results which do not require differentiability, so finite examples are covered. Some
of the results concern the exogenous case and others concern the endogenous case. Section 7
presents three exogenous-case theorems which require differentiability. Section 8 presents four
endogenous-case theorems which require differentiability. Section 9 considers the shape of the
function which relates r to the Principal’s gain. A concave shape implies a proposition about the
negotiation set when the two parties bargain about the size of r. Section 10 provides concluding
remarks about extensions and variations of the model.
5. Some examples
Our model, simple as it is, turns out to have quite diverse results. To illustrate the diversity,
we now consider four examples. In all four of them the effort set and the set of possible values
of the technology parameter t are continua and calculus methods are used to study them. In the
simplest finite example, on the other hand, there would be just two values of t and two values of
effort. We can construct simple finite examples yielding a variety of answers to the questions we
just listed. In each of our three non-finite examples we provide some statements that we shall
subsequently generalize.
5.1 A “Classic monopoly” example where marginal revenue falls and marginal cost
is flat.
For convenient reference we shall call this our Classic monopoly example — or, for brevity,
our Classic example. It is suggested by the introductory monopoly diagram in the typical text,
where marginal revenue drops and marginal cost is flat or rising. We may think of the Principal
as a monopolist who delegates the choice of product quantity to the Agent. Quantity will be our
“effort”. At effort x, price is A − Bx, where A > 0, B > 0. Cost is t · C(x) = tx and revenue is
A
R(x) = Ax − Bx2 . Marginal revenue becomes negative at x = 2B
. To keep price and marginal
revenue positive, our set of possible efforts will be
A
Σ = 0,
.
2B
8
One of them concerns optimal versus “selfish” routing in transportation networks (see Roughgarden (2005)).
Others are found in Nissan, Roughgarden, Tardos, and Vazirani (eds.) (2007).
7
We consider a set Γ of pairs (r, t), where (i) every r ∈ (0, 1) belongs to one and only one pair,
and (ii) for every pair (r, t) ∈ Γ there is an effort x̂(r, t) which belongs to Σ and is the unique
maximizer of r · R(x) − t · C(x). That is the case for
Γ ≡ {(r, t) : 0 < r < 1; 0 < t < Ar}
and
A
t
−
.
2B 2Br
For a given t, the Agent’s best response to the share r — if (r, t) ∈ Γ — is the effort x̂(r, t),
which belongs to Σ.
x̂(r, t) =
In every example that we study we will specify a similar set Γ, where the Agent has a unique
nonnegative maximizing effort for every r ∈ (0, 1). Note that in the Classic example, Γ is the
interior of a triangle. In a diagram with r on the horizontal axis and t on the vertical axis the
triangle has vertices at (0, 0), (1, 0) and (1, A). In other examples Γ might be a rectangle, as in
the example which follows (in 5.2). In still other examples one of the boundaries of Γ might have
curvature. We shall let Γ̃ denote the set of values of t which we consider. Thus
Γ̃ ≡ {t : (r, t) ∈ Γ for some r ∈ (0, 1)}.
In our Classic example, Γ̃ = (0, A).
Now for every (r, t) ∈ Γ consider the derivative of x̂ with respect to r, the derivative with
respect to t, and the cross derivative. They will be denoted by x̂r , x̂t and x̂rt . We have the
following results. Some of them will be generalized to wider classes of examples.
t
• x̂r (r, t) = 2Br
2 , which is positive. For a given t, increasing the share evokes more effort.
9
We shall show that in any example, finite or nonfinite, increasing the share never evokes less
effort.
1
• x̂t (r, t) = − 2Br
, which is negative. When r is fixed and technology improves (when t
drops), the Agent works harder. We shall show 10 that in any example, finite or nonfinite, the
Agent never works less when t drops.
−1
• We have x̂rt (r, t) = 2B
· −1
> 0. So a technology improvement (a drop in t) diminishes
r2
effectiveness (the effort increase evoked by a small rise in r).
• When the Agent uses the best effort x̂(r, t), he receives r · R(x̂(r, t) − t · x̂(r, t). The
derivative of that expression with respect to t is negative.11 So in the exogenous case, technology
improvement is good news for the Agent. We shall provide12 a simple (calculus-free) proof that
this statement holds in any example, finite or nonfinite.
9
In Part (a) of Theorem 1.
In Part (b) of Theorem 1.
11
The derivative is x̂t (r, t) · [rR0 (x̂(r, t)) − t · C 0 (x̂(r, t)] − C(x̂(r, t)). That is negative, since 0 < r < 1 and x̂(r, t)
satisfies the first-order condition 0 = rR0 − tC 0 .
12
In Part (f ) of Theorem 1.
10
8
• We find that
surplus = W (r, t) = R(x̂(r, t)) − t · C(x̂(r, t)) =
1
· [(Ar − t) · (BAr + Bt − 2Brt)].
4B 3 r3
The derivative with respect to t of the expression in square brackets is
−2BAr2 − 2Bt + 4Brt.
Our requirement that t < Ar implies that this is negative.13 For a fixed r < 1, exogenous surplus
rises when technology improves (t drops). We shall see14 that this holds whenever R and C are
differentiable and a first-order condition determines the Agent’s best effort. But it need not hold
in finite examples.
• For all t ∈ Γ̃, we have15 Wt (1, t) < 0. Maximal surplus16 rises when technology improves
(t drops). As we shall see17 , a trivial argument shows that this always holds in both finite and
non-finite examples.
t
• We have Wrt (r, t) = 2Br
2 > 0. So x̂rt (r, t) and Wrt (r, t) have the same sign. That implies
that the exogenous Decentralization Penalty (surplus gap) W (1, t) − W (r, t) and the exogenous
effort gap x̂(1, t)−x̂(r, t) move in the same direction when technology improves, i.e., the exogenous
surplus gap tracks the exogenous effort gap. There are finite examples where that is not the
case. But we shall show18 that if R and C are thrice differentiable then it must be the case,
because, as we shall prove, Wrt · x̂rt ≥ 0.
We now turn to the endogenous case. The Principal chooses a best share r, but excludes
r = 1, which would give the Agent all of the revenue. To study the consequence of choosing
r = 0, we would have to specify how the Agent responds to r = 0. The natural answer is zero
effort, but there would then be some technical difficulties in our analysis of certain examples.
So, as already noted, we confine attention to cases where the Agent’s response to r in the open
interval (0, 1) is a positive effort and we shall suppose that the Principal chooses an r that is
best in (0, 1). We will show19 that under simple conditions (which are satisfied in the Classic
example), the Principal’s gain (1 − r) · (R(x̂(r, t)) is positive for some r ∈ (0, 1) and is concave
on (0, 1). That implies that there is a share in (0, 1) which solves the first-order condition
0=
d
[(1 − r) · R(x̂(r, t)] = −R(x̂(r, t)) + (1 − r) · R0 (x̂(r, t)) · x̂r (r, t)
dr
13
The derivative is negative if Ar2 > t · (2r − 1). That is the case at r = 0 and at r = 1 (since t < A. At all
r ∈ (0, 1) our requirement t < Ar implies that 2Ar, the derivative of the left side of the inequality with respect
to r exceeds 2t, the derivative of the right side. So at all (r, t) ∈ Γ the inequality holds.
14
In Part (b) of Theorem 2.
15
For derivatives, we shall use the symbols Wr , Wt , Wrt , which are analogous to our symbols x̂r , x̂t , x̂rt .
16
In the monopoly setting our general term “surplus” should not be confused with consumers’ surplus. Our
“Surplus” is the same as the monopolists’ profit.
17
In Part (d) of Theorem 1
18
In Theorem 4.
19
In Theorem 8.
9
and maximizes the Principal’s gain on the set (0, 1). That share is our r∗ (t) ∈ (0, 1). In our
Classic example the Principal’s first-order condition turns out to be the cubic equation
0 = A2 r3 + rt2 − 2t2 .
When we graph the implicit function r∗ (t), we obtain the following figure for the case A = 2, B =
3:
FIGURE 1 HERE
The graph reveals that r∗ is increasing in t. The share-choosing Principal becomes less generous when technology improves. But we can establish this fact — for a very wide class of examples
that includes the Classic example — without any graphing. We shall show20 that it must hold
if R00 < 0 and x̂rt < 0 (as in the Classic example). Thus a drop in t makes the Principal less
generous if it makes marginal revenue drop and it also makes effectiveness drop.
Next consider the endogenous effort x̂(r∗ (t), t). Figure 2 shows that in the Classic example
the endogenous effort rises when technology improves (t drops). We shall show21 that this must
happen, for the endogenous case, in every example, finite or nonfinite.
FIGURE 2 HERE
The endogenous decentralization Penalty (surplus gap) is W (1, t) − W (r∗ (t), t). That can be
graphed as a function of t, even without an explicit expression for r∗ (t). We do so in Figure 3
for the Classic example. We see that for large t further technical improvement (a further drop
in t) raises the Penalty, but for small t further improvement lowers it.
FIGURE 3 HERE
We now turn to the tracking question. For the Classic example, Figure 4 shows both the
Penalty (surplus gap) and the effort gap x̂(1, t) − x̂(r∗ (t), t). When t increases each gap first rises
and then falls and for each t the gaps move in the same direction, so we indeed have tracking.
FIGURE 4 HERE
5.2 A “Cubic-revenue” example, where, just as in the Classic example, technology improvement diminishes effectiveness and generosity, but now we do not have
tracking.
In this example:
R(x) = x3 − x2 , C(x) = x, and the set of possible (r, t) pairs is Γ = {(r, t) : r ∈ (0, 1); t ≤ r}
20
21
In Theorem 6.
In Part (j) of Theorem 1.
10
t
Fig. 1 graph of r∗ (t) for the Classic example with A = 2, B = 3
t
Fig. 2 graph of x̂(r ∗ (t), t) for the Classic example with A = 2, B = 3
W (1, t)
W (r∗ (t), t)
t
Figure 3: W (1, t) and W (r∗ (t), t) for the Classic case, with A = 2, B = 3
effort gap
surplus gap
t
Figure 4: The two gaps: surplus gap (W (1, t) − W (r∗ (t), t)) and effort gap (x̂(1, t) − x̂(r∗ (t), t))
for the Classic case, with A = 2, B = 3
and the set of possible values of t is Γ̃ = (0, 1). We find — just as in the Classic example —
that for all r, t) in Γ, effectiveness diminishes when t drops, i.e., x̂rt (r, t) ≥ 0.22 Now consider
r∗ (t), the implicit function which solves the Principal’s first-order equation for t ∈ (0, 1). That
function turns out to satisfy a cumbersome polynomial equation.23 Figure 5 graphs the implicit
function r∗ .
FIGURE 5 HERE
Just as in the Classic example, r∗ rises at every possible value of t (all t ∈ (0, 1)). Finally, we
plot the surplus gap and the effort gap.
FIGURE 6 HERE
We find that for t in the interval (.48, .63), the effort gap rises but the surplus gap falls. Unlike
the Classic example, we do not have tracking.
5.3 A “Price-taker” example where marginal cost rises and marginal revenue is flat
In this example we may think of the Principal as a price-taker (with price equal to one). He
delegates quantity choice to the Agent and the Agent’s cost function is quadratic. “Price-taker”
is a convenient label for this example. The example is defined as follows:
• The set of possible efforts is Σ = IR+ .
• R(x) = x.
22
We have
1/2
1
t
.
x̂(r, t) = √ · 1 −
r
3
Next we obtain
1 1 d
x̂r (r, t) = √ · ·
3 2 dr
"
t
1−
r
1/2 #
−1/2
1
t
t
= √ · 1−
· 2.
r
r
4 3
We then have:
"
#
−1/2
t
d
t
· 2
1−
4 3 · x̂rt (r, t) =
dt
r
r
1/2 t
1
t
=
1−
·
+ 2·
r
r2
r
√
But
d
dt
"
t
1−
r
−1/2 #
−1 d
·
2
dt
"
t
1−
r
−1/2 #!
.
−3/2
−1
t
t
· 1−
=
· 2 < 0.
2
r
r
So we indeed have x̂rt (r, t) ≥ 0.
The equation is
23
0 = 16[r∗ (t)]6 − 12t2 − 27t4 · [r∗ (t)]4 − 4t3 + 108t4 · [r∗ (t)]3 − 162t4 · [r∗ (t)]2 + 108t4 · r∗ (t) − 27t4 .
11
t
Fig. 5. Graph of r∗ (t) in the Cubic-revenue example.
effort gaps
surplus gap
t
Fig. 6 Surplus and effort gaps in the Cubic-revenue example. For t ∈ (.48, .63),
effort gap rises
ses but surplus
falls. but surplus gap falls.
• C(x) = 12 (x − 1)2 .
• The set of possible pairs (r, t) is the rectangle Γ = {(r, t) : 0 < r < 1; 13 < t < 1}.
We find the following:
• Given r, the Agent chooses x̂(r, t) = rt + 1. We then have x̂rt (r, t) =
rises when technology improves (when t drops).
−1
t2
< 0. Effectiveness
2
• When the Agent uses the
best effort x̂(r, t), he receives rt + 1 − r2t . The derivative with respect
to t is is t12 · 12 r2 − r , which is negative, since 0 < r < 1. So, just as in the Classic example
technology improvement is good news for the Agent in the exogenous case. As we already
noted, we will provide a simple (calculus-free) proof that this “good news” statement holds,
for the exogenous case, in any example, finite or nonfinite.
• Exogenous surplus is R(x̂(r, t)) − t · C(x̂(r, t)) =
• Surplus-maximizing effort is
1
t
r
t
+1−
r2
.
2t
+ 1 and maximal surplus is 1 +
1
.
2t
2
• The exogenous surplus gap (the Penalty) is 2t1 − rt + r2t . Its derivative with respect to t is
, which also has a negative derivative. So the
negative. The exogenous effort gap is 1−r
t
exogenous surplus gap tracks the exogenous effort gap, just as in the Classic example. As
already noted, this will be proved to hold,in the exogenous case, for any example.
We now turn to the endogenous case. We find that:
• The solution to the Principal’s first order condition 0 =
0
d
[R(x̂(r, t)) − t · C(x̂(r, t))]
dr
is r∗ (t) =
So we have r∗ (t) < 0 at every possible t. (Recall that t < 1). So, in sharp contrast
to the Classic monopoly example, the Principal becomes more generous when technology
improves.
1−t
.
2
• We find24 that — just as in the exogenous case — a drop in t is good news from the welfare
point of view. This must be the case whenever — as in the Price-taker example and the
Rising Marginals example which we consider next — the Principal becomes more generous
(or stays just as generous) when t drops.25
• The endogenous Penalty (the endogenous surplus gap) is 14 + 8t1 + 8t . Its derivative with
respect to t is 8t12 · (t2 − 1), which is negative, since t < 1. The Penalty rises when
technology improves. Again, note the contrast with the Classic example, where the Penalty
drops when technology improves, once t has dropped below a critical value.
24
In Part (l) of Theorem 1
We have
0
d
[R(x̂(r∗ (t), t)) − t · C(x̂(r∗ (t), t))] = [x̂r · r∗ + x̂t ] · (R0 − tC 0 ) − C.
dt
0
0
We have R − tC > 0 (because of the first-order condition rR0 − tC 0 = 0, where 0 < r < 1). Since x̂t < 0 and
0
r∗ ≤ 0, we conclude that the derivative is negative, so we indeed have “good news” from the welfare point of
view.
25
12
• The endogenous effort gap is
1+t
t
−
r∗ (t)+t
t
=
1
2t
+ 12 . That also has a negative derivative.
• So the endogenous effort gap tracks the exogenous surplus gap. But that is NOT implied, as
0
we shall see, by the fact that x̂rt < 0 and r∗ (t) < 0.
5.3 An example where marginal revenue rises but marginal cost rises faster.
It will be convenient to call this the “Rising Marginals” example. We have:
• The set of possible efforts is Σ = IR+ .
• R(x) = xa , C(x) = xb , 0 < a < b.
• The set of possible pairs (r, t) is Γ = {(r, t) : 0 < r < 1; t > 0}.
We obtain the following:
1
tb a−b
• x̂(r, t) = ra
.
1/(a−b) 1/(b−a)−1
1
• x̂rt (r, t) = a−b
· t1/(a−b)−1 · ab
·r
. That is negative, since a < b. When
technology improves effectiveness increases.
• In the endogenous case the Principal chooses the share r∗ (t) =
generosity remains unchanged when technology changes.26
a
.
a+1
The Principal’s
• Just as in the Price-taker example, Improvement in technology is good news for the share0
choosing Principal. That is the case because r∗ = 0.
• Even though we have an explicit expression for r∗ , computing the derivative of endogenous
effort gap (Penalty) with respect to t and the derivative of endogenous surplus gap with respect
to t is cumbersome. It turns out that both are negative. So the endogenous surplus gap tracks the
endogenous effort gap. This is not true in the Classic example. While it is true in the Price-taker
26
The first-order condition satisfied by r∗ can be written
r =1−
R(x̂(r, t))
.
· x̂r (r, t)
R0 (x̂(r, t))
In the example we obtain:
x
R(x)
= ,
0
R (x)
a
x̂r =
tb
a
1/(a−b)
·
1
· r1−1/(b−a) ,
b−a
x̂
= r.
x̂r
So the first-order condition is
r =1−
That is solved by r∗ =
1 x̂
r
·
=1− .
a x̂r
a
a
a+1 .
13
0
example, that does not follow from the signs of x̂rt and r∗ in that example. In contrast, we shall
show27 that, in the endogenous case, the surplus gap tracks the effort gap whenever (as in the
0
Rising Marginals example) x̂rt < 0 and r∗ ≥ 0.
6. Basic results that do not require differentiability.
The following twelve-part theorem applies to all examples, finite and nonfinite. An example
is defined by a set Σ of possible positive efforts, the functions R and C, and a set Γ of possible
pairs (r, t). Recall that for every r ∈ (0, 1), Γ contains some pair (r, t). The set of values of t such
that (r, t) ∈ Γ for some r is again denoted Γ̃. Recall that x̂(r, t) denotes the effort which is the
smallest maximizer of rR(x)−tC(x) on Σ, that W (r, t) denotes the surplus R(x̂(r, t))−tC(x̂(r, t),
and that r∗ (t) denotes the smallest maximizer of (1 − r) · R(x̂(r, t)) on (0, 1).
Parts (a), (b) say that in the exogenous case the Agent never works less when the share r
rises and when t drops (technology improves). Part (c) says that the surplus-maximizing effort
cannot fall when t drops. Part (d) says that maximal surplus must rise when t drops. Parts
(e),(f ) say that in the exogenous case a drop in t is never bad news for the Principal and never
bad news for the Agent, respectively. Part (g) says that the drop in t must be good news from
the welfare point of view. Part (h) says that in the exogenous case a rise in the share must be
good news for the Agent. Parts (i), (j), (k), (l) concern the endogenous case. Part (i) says
that the ratio of the Principal’s chosen share to the technology parameter t cannot fall when t
drops. But, as we have already seen in the examples, the chosen share itself may rise or fall or
stay the same. Nevertheless Part (j) says that in the endogenous case the Agent never works
less hard when t drops. Part (k) says that in the endogenous case a drop in t is never bad news
for the Principal. Part (l) says that in the endogenous case a drop in t must be good news from
the welfare point of view.
Theorem 1
Let R and C be strictly increasing on Σ. Then:
(a) x̂(rH , t) ≥ x̂(rL , t) whenever (rL , t) ∈ Γ, (rH , t) ∈ Γ, and 0 < rL < rH < 1.
(b) x̂(r, tH ) ≥ x̂(r, tL ) whenever (r, tL ) ∈ Γ, (r, tH ) ∈ Γ, and 0 < tL < tH .
(c) x̂(1, tL ) ≥ x̂(1, tH ) whenever tL , tH ∈ Γ̃, and 0 < tL < tH .
(d) W (1, tL ) > W (1, tH ) whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
(e) (1 − r) · R(x̂(r, tL )) ≥ (1 − r) · R(x̂(r, tH )) whenever (r, tL ) ∈ Γ, (r, tH ) ∈ Γ, and 0 < tL < tH .
(f ) rR(x̂(r, tL )) − tL C(x̂(r, tL )) > rR(x̂(r, tH )) − tH C(x̂(r, tH )) whenever (r, tL ) ∈ Γ, (r, tH ) ∈ Γ,
and 0 < tL < tH .
(g) W (r, tL ) > W (r, tH ) whenever (r, tL ) ∈ Γ, (r, tH ) ∈ Γ, and 0 < tL < tH .
27
In Part (a) of Theorem 5.
14
(h) rH ·R(x̂(rH , t))−tC(x̂(rH , t)) > rL ·R(x̂(rL , t)−tC(x̂(rL , t)) whenever (r, tL ) ∈ Γ, (r, tH ) ∈ Γ,
and 0 < rL < rH .
(i)
r∗ (tL )
tL
≥
r∗ (tH )
tH
whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
(j) x̂(r∗ (tL ), tL ) ≥ x̂(r∗ (tH ), tH ) whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
(k) (1 − r∗ (tL )) · R(x̂(r∗ (tL ), tL ) ≥ (1 − r∗ (tH )) · R(x̂(r∗ (tH ), tH ) whenever tL , tH ∈ Γ̃ and
0 < tL < tH .
(l) W (r∗ (tL ), tL ) > W (r∗ (tH ), tH ) whenever tL ∈ Γ̃, tH ∈ Γ̃, and 0 < tL < tH .
In proving Parts (e),(f ),(h) we use the simple observation that when t drops from tH to tL
and when r rises from rL to rH , the Agent could continue to use the previous effort. Similarly,
in the endogenous Part (k), the Principal could continue to use the share r∗ (tH ). The proof
of Part (g) uses part (b) and in a similar way the proof of Part (l) uses part (j). In proving
Parts (a),(b),(c),(d),(i),(j),28 which concern the maximizer x̂ and the maximizer r∗ , we use a
standard proposition from monotone comparative statics29 :
If a function h(u, v) displays strictly increasing differences, and if uH maximizes h(u, vH )
while uL maximizes h(u, vL ), then uH ≥ uL if vH > vL .
The proof of Theorem 1, like all the subsequent proofs, is found in the Appendix.
Note that Part (k) of Theorem 1 tells us that in the endogenous case technical improvement
must be good news for the Principal. But for the Agent, the situation is different. Figure 7 is a
graph of the Agent’s endogenous-case net earnings r∗ (t) · R(x̂(r∗ (t)), t) − t · C(x̂(r∗ (t)), t) in the
Classic example. Once t has dropped to a critical value that is close to 0.5, a further drop is Bad
news for the Agent. To put it in a crude but colorful way: in the endogenous case, the Principal
is never the enemy of technical progress but the Agent might be.
FIGURE 7 HERE
7. Three exogenous-case theorems which require differentiability.
7.1. Welfare (surplus) increases when the share increases and when technology
improves (t drops).
Theorem 2
Suppose R and C are differentiable and x̂(r, t) is the solution to the first-order equation rR0 (x) =
tC (x). Then (a) W (rH , t) > W (rL , t) whenever 0 < rL < rH , and (b) W (r, tL ) > W (r, tH )
whenever 0 < tL < tH .
0
28
Note that in (a) and (b) the inequalities are weak. In the differentiable case appropriate conditions on C and
R make those inequalities strict.
29
See, for example, Sundaram (1996)
15
t
Fig. 57 The Agent’s net earnings for the Classic example with A = 2, B = 3
7.2 Effectiveness and the effort gap move in the same direction when t changes.
Theorem 3
Let Γ be an open set in IR+ . Suppose that the functions R and C are thrice differentiable. Suppose
that the Agent’s best-effort function x̂ satisfies the first-order condition rR(x̂(r, t))−tC 0 (x̂(r, t)) = 0.
Then x̂(r, t) ≥ 0 (< 0) at every (r, t) ∈ Γ if and only if
d
[x̂(1, t) − x̂(r, t)] ≥ 0 (< 0) at every (r, t) ∈ Γ.
dt
Note that the pair (r∗ (t), t) belongs to Γ, so the theorem applies, in particular, to x̂rt (r∗ (t), t)
and the endogenous effort gap x̂(1, t) − x̂(r∗ (t), t).
Theorem 3 is directly implied by another standard result in monotone comparative statics.
That result is as follows:
Let h be a function of the variables u, v, let h be twice differentiable at all (u, v) in
an open set Γ ⊆ IR2 , and let huv denote the cross-partial. Then h displays increasing
differences on Γ if and only if hu,v (ū, v̄) ≥ 0 at all (ū, v̄) ∈ Γ.
If we now apply this standard result to the function x̂t (r, x), we see that effectiveness and the
effort gap indeed move in the same direction when t changes. In the Appendix, however, we
provide an alternative self-contained proof that uses the Mean Value Theorem and does not
appeal directly to the standard result.
7.3. A theorem about the effort gap and the surplus gap.
Theorem 4, which now follows, says that in the exogenous case we have “tracking”: for any
fixed r the surplus gap W (1, t) − W (r, t) and the effort gap x̂(1, t) − x̂(r, t) must move in the
same direction when t changes. That is not true in general for the endogenous case, but it
becomes true under appropriate conditions on the direction in which share-increase effectiveness
and Principal’s generosity move. We supply such conditions in Theorem 5. As we shall see,
the proof of the endogenous-case Theorem 5 relies heavily on the exogenous-case Theorem 4. In
Theorem 4 the effort set Σ is an open interval (0, J). The set Γ of possible (r, t) pairs is an open
set in IR+ and has our standard property: for every r ∈ (0, 1), Γ contains some pair (r, t). As
before, Γ̃ denotes the set {t : (r, t) ∈ Γ for some r}.
Theorem 4
Let Γ be an open set in IR2 . Suppose that on the effort set (0, J), with J > 0, the functions R
and C are positive-valued and thrice differentiable. Suppose R0 > 0, C 0 > 0. Suppose that the
Agent’s best-effort function x̂ : Γ → (0, J) satisfies the condition rR0 (x̂(r, t)) − tC 0 (x̂(r, t)) = 0
for every (r, t) ∈ Γ, and that the surplus-maximizing-effort function x̂(1, · ) : Γ̃ → (0, J) satisfies
R0 (x̂(1, t)) − tC 0 (x̂(1, t)) = 0 for every t ∈ Γ̃.
Then we have
16
(+)
For any r ∈ (0, 1), x̂(1, tL ) − x̂(r, tL ) > x̂(1, tH ) − x̂(r, tH ) ( x̂(1, tL ) −
x̂(r, tL ) < x̂(1, tH ) − x̂(r, tH ) ) whenever tL , tH ∈ Γ and 0 < tL < tH
if and only if we also have

∈
(0, 1), W (1, tL ) − W (r, tL )
>
W (1, tH ) −
For any r
(++)
W (r, tH ) ( W (1, tL ) − W (r, tL ) < W (1, tH ) − W (r, tH ) ) when
ever tL , tH ∈ Γ̃ and 0 < tL < tH .
[The surplus gap (Penalty) is increasing (decreasing) in t if and only if the effort gap is increasing
(decreasing) in t].
Straightforward calculation yields the following Corollary.
Corollary
Suppose the Theorem’s conditions on the functions R, C, the best-effort function, and the surplusmaximizing effort function are all met. Then
• the Decentralization Penalty (surplus gap) is decreasing in t (so the Penalty grows when technology
improves) if at every effort x ∈ (0, J) we have R00 (x) > 0, R000 (x) = C 000 (x) = 0.
• the Decentralization Penalty (surplus gap) is increasing in t (so the Penalty shrinks when technology
improves) if at every effort x ∈ (0, J) we have R00 (x) < 0, C 00 (x) = 0, R000 (x) ≤ 0.
Even though we are in the relatively straightforward exogenous case, the Corollary’s sufficient
conditions for the Penalty to grow (shrink) when technology improves are simple but quite
restrictive. When we turn to the endogenous case, we find no similarly simple conditions on R
and C which tell us, all by themselves, the direction in which the Penalty moves when technology
improves.
8. Endogenous-case results which require differentiability.
We have seen that in the endogenous case there are examples where the Decentralization
Penalty (surplus gap) rises when tech ology improves and there are examples where it falls. There
are examples where we have “tracking” (surplus and effort gaps move in the same direction when
t changes), but there other examples where that is not true.
Can we categorize examples so that there is some order in this diversity? The key to the
desired order turns out to be effectiveness and Principal’s generosity. “Effectiveness grows when
technology improves (t drops)” means that the cross partial x̂rt (r, t) is negative at every (r, t)
in Γ. “Effectiveness shrinks when technology improves” mans that the cross partial is positive.
0
0
“Generosity increases (decreases) when technology improves” means that r∗ (t) < 0 (r∗ (t) > 0).
0
“Generosity stays the same” (as in the Rising Marginals example) means that r∗ (t) = 0.
8.1 Two endogenous-case theorems which concern the effort and surplus gaps and
require differentiability.
We now develop Theorems 5 and 6, which concern the tracking question. The following fourbox table will serve as a guide to the theorems and their relation to the examples. The table
concerns what we will call Interior Examples. An example is a quadruple (Σ, R, C, Γ). It is an
Interior Example if
17
• Σ and Γ are open sets. (In each set, every point has a neigborhood which is a subset of that
set).
• R, C are thrice differentiable on Σ and R0 > 0, C 0 > 0.
• For every (r, t) ∈ Γ, there exists a unique effort x̂(r, t) ∈ Σ which satisfies the first-order
condition 0 = rR0 (x) − tC 0 (x) and maximizes rR(x) − tC(x) on Σ.
• For every t ∈ Γ̃, there exists a unique share r∗ (t) ∈ (0, 1) which satisfies the first-order
d
condition 0 = dr
[(1 − r) · R(x̂(r, t))] and maximizes (1 − r) · R(x̂(r, t)) on (0, 1).
The Classic example, the Price-taker example, and the Rising Marginals example satisfy these
conditions.
18
THE EFFECT OF IMPROVED TECHNOLOGY (A DROP IN t) IN FOUR GROUPS OF INTERIOR EXAMPLES
WHEN
t
DROPS,
EFFECTIVE-
NESS OF A SHARE INCREASE
WHEN
t
DROPS,
EFFEC-
TIVENESS OF A SHARE IN-
AND
CREASE RISES AND HENCE
HENCE THE EFFORT GAP ALSO
THE EFFORT GAP ALSO
FALLS
RISES (SEE THEOREM 3).
FALLS OR STAYS THE SAME
OR
STAYS
THE
SAME
(SEE THEOREM 3).
x̂rt ≥ 0 and
d
[x̂(1, t) − x̂(r, t)] ≥ 0
dt
1
WHEN t DROPS, PRINCI-
x̂rt < 0 and
d
[x̂(1, t) − x̂(r, t)] < 0
dt
SEE “CLASSIC” EXAMPLE.
2
SEE “RISING
PAL BECOMES MORE GEN-
EFFORT AND SURPLUS GAPS MOVE IN THE
MARGINALS”
EROUS
SAME DIRECTION IN SOME EXAMPLES AND
EFFORT AND SURPLUS GAPS MUST
OR
GENEROSITY
STAYS THE SAME.
IN OPPOSITE DIRECTIONS IN OTHERS. SEE
FIGURES 4 AND 5, WHICH CONCERN
0
r∗ ≥ 0
VARIATIONS OF THE “CLASSIC” EX-
EXAMPLE.
MOVE IN THE SAME DIRECTION
(Theorem
(a)).
5,
Part
AMPLE.
3
WHEN t DROPS, PRINCI-
SEE “EXPLODING
4
SEE “PRICE-
PAL BECOMES LESS GEN-
MARGINALS” EXAMPLE.
EROUS.
AND SURPLUS GAPS MUST MOVE IN THE
EXAMPLE EFFORT AND SURPLUS
(Theorem 5,
GAPS MOVE IN THE SAME DIREC-
AN EXAMPLE WITH
TION, BUT EXAMPLES CAN BE CON-
SAME DIRECTION
Part (b)).
0
r∗ < 0
R
00
<
0, x̂t
<
EFFORT
0 AND x̂r
CANNOT BE IN THIS BOX
>
0
(Theo-
TAKER” EXAMPLE.
IN THAT
STRUCTED WHERE THEY MOVE IN
OPPOSITE DIRECTIONS.
rem 6).
The example in Box 3, which we call the “Exploding Marginals” example, has functions R
and C such that C 00 > R00 and both C 00 and R00 rise extremely rapidly. It appears difficult to
construct Box-3 examples where that is not the case.30 The Exploding Marginals example is as
follows:
• Σ = (0, 1).
• Γ = {(r, t) : 0 < r < 1; rt ∈ (e, ee )} (e is the base of the natural logarithms).
2
• R(x) = ex .
30
One can prove, for example, that if R = 12 x2 , so that R00 = 1, then we cannot be in Box 3.
19
• C(x) =
R x ep p 2e · e · p dp.
0
The proof that the Exploding Marginals example indeed lies in Box 3 is provided in the Appendix.
We now have Theorem 5, a two-part theorem, which concerns Box 2 and Box 3.
Theorem 5
Consider an interior example (Σ, Γ, R, C).
(a). Suppose the following holds:
0
for every t ∈ Γ̃ we have r∗ (t) ≥ 0 and for every (r, t) ∈ Γ we have x̂rt (r, t) < 0.
Then for every (r, t) ∈ Γ we have
∂
∂
∗
∗
x̂(1, t) − x̂(r (t), t) ·
W (1, t) − W (r (t), t) ≥ 0.
∂t
∂t
(The surplus gap and the effort gap move in the same direction when t changes).
(b). Suppose the following holds:
0
for every t ∈ Γ̃ we have r∗ (t) < 0 and for every (r, t) ∈ Γ we have x̂rt (r, t) ≥ 0.
Then for every (r, t) ∈ Γ we have
∂
∂
∗
∗
x̂(1, t) − x̂(r (t), t) ·
W (1, t) − W (r (t), t) ≥ 0.
∂t
∂t
(The surplus gap and the effort gap move in the same direction when t changes).
The next theorem does not directly concern the two gaps. But it implies that if marginal
revenue is decreasing or constant (R00 ≤ 0) in an interior example and the Principal has a unique
best share, then the example cannot be in Box 3.
Theorem 6
Suppose that in the interior example (Σ, Γ, R, C) we have:
• R00 (x) ≤ 0 at every x ∈ Σ.
• x̂rt (r, t) ≥ 0, x̂t (r, t) < 0 and x̂r (r, t) > 0 at every (r, t) ∈ Γ.
• r∗ (t) is the unique maximizer of (1 − r) · R(x̂(r, t)) on (0, 1),
20
0
Then r∗ (t) ≥ 0 for all t ∈ Γ.
It is difficult to give a clear intuition for Theorems 4 and 5. That is a little easier for Theorem
6, which says that if marginal revenue is decreasing, and effectiveness drops when technology
0
improves, then when technology improves, the Principal does not become more generous (r∗ (t) ≥
0), i.e., we cannot be in Box 3. Intuitively one might say: when t drops, increasing the share
above its previous level would damage the Principal, because the extra revenue due to extra effort
has dropped (marginal revenue has declined) and at the same time the extra effort evoked by a
share increase has dropped as well.
8.2 The effect of technical improvement on the Decentralization Penalty (surplus
gap).
We can now say something about our original puzzle: when does technical improvement lower
the Penalty and when does it raise the Penalty? Theorems 4 and 5, together with Part (a) of
Theorem 1, easily yield the following Theorem.
Theorem 7
Consider any interior example. If, in that example, technical improvement makes the Principal less
generous or keeps his generosity unchanged, while at the same time it decreases the effectiveness of
0
a share increase (r ∗ (t) ≥ 0 and x̂rt (r, t) < 0), then the improvement raises the Penalty or keeps it
unchanged. If technical improvement makes the Principal more generous, while at the same time it
0
decreases the effectiveness of a share increase or leaves it unchanged (r ∗ (t) < 0 and x̂rt (r, t) ≥ 0),
then the improvement lowers the Penalty or keeps it unchanged.
Unfortunately there are no simple conditions on R and C, similar to those in the Corollary
to the exogenous-case Theorem 3, which imply, all by themselves, that the Penalty rises (falls)
when technology improves.
9. Finding the Principal’s best share for a given t: when is the Principal’s gain a
concave function of the share?
The function r∗ (t) may be increasing on the set Γ̃ of possible values of t. It may also be
decreasing or constant. We have discussed the implications of each case. But we have not yet
studied, in a general way, the shape of the Principal’s gain as a function of r ∈ (0, 1) when t
is fixed. The gain for fixed t is (1 − r) · R(x̂(r, t)). The graph of the non-negative values of
(1 − r) · R(x̂(r, t)), with r on the horizontal axis, starts at zero and ends at zero. The graph
coincides with the Principal’s gain curve except at r = 0 and r = 1, since the Principal confines
attention to the open interval (0, 1). It would be particularly helpful if the gain curve rises and
then falls, achieving its maximum height at r∗ (t). More generally, the curve could rise until
r = r∗ (t) and could then be flat for an interval before descending. Let us call such a gain
curve single-peaked. As long as the gain is positive at some r ∈ (0, 1) the curve is single-peaked
if it is concave on (0, 1). The following theorem provides conditions under which the gain is
indeed concave. The theorem has two parts. The first part does not require differentiability with
respect to r, but the second part does. Informally, the second part says that we have concavity
21
if marginal revenue drops (R00 < 0) and in addition the effectiveness of a share increase drops
when the share increases (x̂rr < 0).
Theorem 8
(a) If, for a fixed t, R(x̂(r, t) is concave on (0, 1), then the Principal’s gain (1 − r) · R(x̂(r, t) is also
concave on (0, 1).
(b) Consider an interior example (Σ, Γ, R, C) where Σ = (0, J), with J > 0. Then R is concave
on (0, J) if for all x ∈ (0, J) we have R00 (x) < 0, and for all (r, t) ∈ Γ we have x̂rr (r, t) < 0. If
R00 (x) < 0, then a sufficient condition for x̂rr < 0 is
r · R000 (x) − t · C 000 (x) ≤ 0.
Suppose that the Principal’s gain curve is indeed single-peaked, and suppose that the share
the Principal uses is determined by bargaining between the Principal and the Agent. In the
interval between zero and the peak (the interval (0, r∗ (t)]), the Agent strictly benefits from a
rise in r, as we established in Part ((a) of Theorem 1. The Principal prefers a higher r as well
(since the gain curve is rising in the interval). But in the interval between the peak and zero (the
interval (r∗ (t), 1)), where the gain curve is falling, the Agent prefers a higher r and the Principal
prefers a lower r. (If r∗ (t) is followed by a flat interval, then the Principal is indifferent between
shares in the flat interval, but is damaged by share increases beyond the flat interval). So the
negotiation set, where bargaining occurs, is the interval (r∗ (t), 1). Now consider the outcome
of the bargaining from the welfare point of view. Assume that R and C are differentiable and
that x̂(r, t) is the solution to the first-order equation 0 = r · R(x) − t · C(x). Then we know,
from Theorem 2, that for a fixed t, welfare increases when the exogenous share increases. So,
informally speaking, increasing the Agent’s bargaining strength increases welfare. A formal model
of the bargaining process is needed to make this statement precise. Such a model might also
reveal the welfare implications of the fact that if r∗ is increasing in t, then the negotiation interval
(0, r∗ (t)) shrinks when t drops.
10. Concluding remarks.
Recall our central question: does technical improvement strengthen or weaken the case for
full Agent autonomy? One might have reasonably hoped for a straightforward answer since
our revenue-sharing Principal/Agent model is so simple. Specifically one might have hoped
that a natural condition like rising marginal cost and falling marginal revenue unambiguously
implies that the Decentralization Penalty rises (or falls) when technology improves. Instead we
have found that there is no easy answer to our central question. On the other hand, we have
found a rich array of other results. One of them is that in both the exogenous case and the
endogenous case, an advance in technology increases welfare. Another is that an advance in
technology causes the Agent to work harder. That is obvious in the exogenous case, since the
Agent benefits from the advance even if he continues to use his previous effort. It is not obvious
in the endogenous case. Other interesting results for the challenging endogenous case concern
the tracking question. The Principal wants to know whether a technical advance strengthens
22
his preference for Agent autonomy, or whether, instead, it has raised the Penalty so high that
perfect monitoring and policing has now become attractive. If the effort gap always moves in
the same direction as the surplus gap (the Penalty), then it suffices to observe (but not police)
the Agent’s effort before and after the advance. To verify the tracking property we look at the
change in the Principal’s generosity and at effectiveness (the increase in Agent effort for a small
share increase). Effectiveness is easy to compute in many examples.
Can we obtain an easier answer to our central question if we vary or complicate the model?
There are many ways to do so. Here are a few of them.
• Change the definition of “Decentralization Penalty”. As we have already ∗noted,
W (x̂(r,t))
(t),t))
one could let the Penalty be the ratio W
(or, in the endogenous case, WW(x̂(r
),
(x̂(1,t))
(x̂(1,t))
rather than the difference, which we have been considering. Our central question becomes
technically harder and preliminary exercises suggest that it again has no simple answer.
It appears, again, that there are no simple conditions on R and C implying that the
redefined Penalty rises or falls when t drops. In the Classic example, for instance, we again
find (in the endogenous case) that when t rises, the redefined Penalty first rises and then
falls. Moreover, we can easily find propositions which are reversed when we move from the
difference definition of Penalty to the ratio definition.31
31
In the exogenous case, with r fixed, consider the derivative of Penalty with respect to t. For the difference
definition we have
d
[W (1, t) − W (r, t)] = Wt (1, t) − Wt (r, t),
dt
which is negative if
(+)
Wt (1, t) < Wt (r, t).
For the ratio definition we have
2
1
d W (r, t)
=
· [W (1, t) · Wt (r, t) − W (r, t) · Wt (1, t)] .
dt W (1, t)
W (1, t)
But that is positive if (+) holds, since we know that W (1, t) ≥ W (r, t). If (+) fails to hold, then whether the two
derivatives have opposite signs remains open. We have to look at the functions R and C.
Now consider the endogenous case. We have
0
d
[W (1, t) − W (r∗ (t), t] = Wt (1, t) − [Wr · r∗ + Wt ].
dt
(Here Wr , Wt are abbreviations for Wr (r∗ (t), t) and Wt (r∗ (t), t)). We know from Part (d) of Theorem 1 that
Wt ≤ 0. Hence the derivative for the difference definition is negative if
(++)
0
Wr · r∗ + Wt > 0.
On the other hand, for the ratio definition we have
2 h
i
1
d W (r∗ (t), t)
∗0
∗
=
· Wr · r + Wt · W (1, t) − W (r (t), t) · Wt (1, t) .
dt
W (1, t)
W (1, t)
Since Wt ≤ 0, the whole expression is positive if (++) holds.
23
• Introduce uncertainty about revenue for a given effort. Here we rejoin the standard
moral-hazard literature briefly reviewed in our Related Literature section. Revenue depends
on effort and on a random variable whose distribution is known to both parties.
• Introduce uncertainty about the technology parameter t. This variation is trivial if
Principal and Agent are risk-neutral and if W becomes an expected value. Simply replace t,
by its expected value, say t̄. If we abandon risk neutrality, then it is conceivable that there
are specific probability distributions of t, and specific utility functions for the Principal and
the Agent under which we get a simple answer to our central question.
• Replace linear sharing by a more complicated reward scheme. In the nonfinite case,
let the Principal offer the agent a reward of ρ(R) if revenue turns out to be R. Can we find
a non-linear function ρ for which we get a simple answer to our central question? If so, does
the Agent find the function ρ acceptable, and does the Principal prefer it to other functions
which the Agent accepts? In the finite case, can we find a reward for every possible revenue
such that the vector of rewards is accepted by the Agent, is preferred by the Principal, and
implies that the Penalty falls (rises) when t drops?
• Introduce informational asymmetry. Let t be a random variable observed by only one of
the two parties. The other knows the probability distribution of t. If it is the Agent who
observes t, then his best effort x̂(r, t) is a random variable for the Principal and so is the
revenue R(x̂(r, t)). Even if the Principal is risk-neutral and we retain our linear sharing
scheme, it appears difficult to find simple conditions on R, C which imply that the Penalty
rises (or falls) when technology improves. That is especially true for the endogenous case.
• Many agents. In the easiest case there are two Agents, the parameter t is known to all three
parties, both Agents have the same function C, and we retain linear sharing. The realized
revenue R is a function of t and the Agents’ efforts. The Principal chooses two shares whose
sum must lie between zero and one. Then for every given t we have a three-player game.
Each Agent chooses an effort xti and the Principal chooses the two shares, r1t , r2t . Agent
i’s payoff is rit · R(xt1 , xt2 ) − t · C(xti ) and the Principal’s payoff is (1 − r1t − r2t ) · R(xt1 , xt2 ).
Suppose that for every t the game has a pure-strategy equilibrium where the Principal
chooses (r̃1t , r̃2t ) and Agent i chooses the effort x̃ti , and suppose that for a given t, surplus is
maximized by the efforts (x̄t1 , x̄t2 ). The Decentralization Penalty at the equilibrium is
R(x̄t1 , x̄t2 ) − t · C(x̄t1 ) − t · C(x̄t2 ) − R(x̃t1 , x̃t2 ) − t · C(x̃t1 ) − t · C(x̃t2 ) .
When t drops, does the equilibrium Penalty rise or fall?
It was natural to start with our stripped-down model, where we already saw the unexpected
challenges posed by our central question. The question of the effect of improved technology on
the merits of alternative modes of organizing is well motivated but has seldom been the focus
of previous research. The variations and extensions that we have noted, and numerous others,
merit further attention.
24
APPENDIX
Proof of Theorem 1
In proving Parts (a), (b), (c), (d), (i), (j) we shall use a standard proposition from monotone
comparative statics.
Consider sets U ∈ IR, V ∈ IR and a function h : U × V → IR. The two arguments of h are
denoted u, v. The function h displays strictly increasing differences in the variables u, v if
h(uH , vH ) − h(uL , vH ) > h(uH , vL ) − h(uL , vL )
whenever uH , uL ∈ U , vH , vL ∈ V , uH > uL , and vH > vL . As noted in the text, we use the
following proposition from monotone comparative statics:

Suppose that for every v ∈ V , the problem







maximize h(u, v) subject to u ∈ U

(*)


has at least one solution. Suppose also that h satisfies strictly increasing differences in u, v.



Consider vH , vL ∈ V with vH > vL . Let uH be a maximizer of h(u, vH ) on U and let uL be


a maximizer of h(u, vL ) on U . Then uH ≥ uL .
Note the following:
(α) If h takes the form h(u, v) = f (u, v) + g(u), then h displays strictly increasing differences
in u, v if and only if f displays strictly increasing differences in u, v.
(β) If h takes the form h(u, v) = f (u) · g(v) and f is strictly increasing while g is nondecreasing,
then h displays strictly increasing differences in u, v.
(γ) If h takes the form h(u, v) = u · g(v) and g is nondecreasing, then h displays strictly
increasing differences in u, v.
We begin with (a), (b), (c), (d), (i), (j), which use proposition (*). We then proceed to (e),
(f), (g), (h), which do not use (*) and concern the exogenous case. We conclude with (k), (l),
which do not use (*) and concern the endogenous case.
Proof of Part (a)
By (α), the function r · R(x) − tC(x), where t is fixed, displays strictly increasing differences
in r, x if r · R(x) displays strictly increasing differences in r, x. But, in view of (γ), that is the
case, since R is nondecreasing. Since, for fixed t, the effort x̂(r, t) maximizes r · R(x) − tC(x) on
the effort set Σ, Proposition (*) implies x̂(rH , t) ≥ x̂(rL , t), as (a) asserts.
Proof of Part (b)
By (α), the function r · R(x) − tC(x), where r ∈ (0, 1) is fixed, displays strictly increasing
differences in −t, x if −t · C(x) displays strictly increasing differences in −t, x. By (γ), that is
25
the case, since C is nondecreasing. Since, for fixed r, the effort x̂(r, t) maximizes r · R(x) − tC(x)
on Σ, Proposition (*) implies x̂(r, tL ) ≥ x̂(r, tH ), as (b) asserts.
Proof of Part (c)
By (α) and (γ), R(x) − t · C(x) displays strictly in −t, x. The effort x̂(1, t) is a maximizer of
R(x) − t · x. Hence, by Proposition (*), x̂(1, tL ) ≥ x̂(1, tH ), as Part (c) asserts.
Proof of Part (d)
Recall that W (1, t) is the maximal surplus for a given t. Using (c), and the fact that R and
C are nondecreasing, we have
W (1, tL ) = R(x̂(1, tL )) − tL · C(x̂(1, tL )) ≥ R(x̂(1, tH )) − tL · C(x̂(1, tH ))
≥ R(x̂(1, tH )) − tH · C(x̂(1, tH )) = W (1, tH ),
as (d) asserts.
Proof of Part (i)
This part concerns the ratio rt , which we shall denote ρ(t). For fixed t, the set of possible
ratios is 0, 1t . Given a ratio ρ and the share r = t · ρ, the Agent’s chosen effort x̂(r, t) is the
smallest maximizer of t · ρ · R(x) − tC(x) on the effort set Σ. Let φ(ρ) be a new symbol, denoting
the effort the Agent chooses when the ratio is ρ. Thus φ(ρ(t)) = x̂(tρ(t), t).
We now claim that
(1)
φ(ρH ) ≥ φ(ρL ) whenever 0 < ρL < ρH .
To see this, suppose that for a fixed r we have
r
r
ρH = ∗ > ∗∗ = ρL .
t
t
Then t∗ < t∗∗ . So, by Part (b), the Agent works at least as hard at ρH as at ρL , i.e., φ(ρH ) ≥
φ(ρL ).
Using the symbols
ρ, φ, we can restate the Principal’s choice for a given t. He chooses the
r∗ (t)
∗
ratio ρ (t) = t , where
)
(
ρ∗ (t) = min arg max M (ρ, −t)
ρ∈(0, 1t ]
,
and
M (ρ, −t) ≡ (1 − tρ) · R(φ(ρ)) = R(φ(ρ)) − t · ρ · R(φ(ρ)).
In view of (α), the function M has strictly increasing differences in ρ, −t if the function −t · ρ ·
R(φ(ρ)) has strictly increasing differences in ρ, −t. But that is the case (using (γ)) since R is
nondecreasing, which implies (using (1)) that R(φ( · )) is also nondecreasing. Since ρ∗ (t) is a
maximizer of M (ρ, −t), Proposition (*) then implies that
(2)
r∗ (tH )
r∗ (tL )
∗
∗
= ρ (tL ) ≥ ρ (tH ) =
whenever 0 < tL < tH ,
tL
tH
26
as (i) asserts.
Proof of Part (j)
We use the terminology in the proof of Part (i). Since φ
(1), x̂(r∗ (tL ), tL ) ≥ x̂(r∗ (tH ), tH ), as (j) asserts.
r∗ (t)
t
= x̂(r∗ (t), t), we have, using
Proof of Part (e)
This follows immediately from (b) and the fact that R is nondecreasing.
Proof of Part (f)
Since x̂(r, t) is a maximizer of R(x) − t · C(x), we have
rR(x̂(r, tL )) − tL · C(x̂(r, tL )) ≥ rR(x̂(r, tH )) − tL · C(x̂(r, tH )) > rR(x̂(r, tH )) − tL · C(x̂(r, tH )).
That implies (f).
Proof of Part (g)
Part (g) says:
W (r, tL ) > W (r, tH ) whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
The effort x̂(r, tL ) is a maximizer of rR(x) − tH · C(x). Hence
r · R(x̂(r, tL )) − tL · C(x̂(r, tL )) ≥ r · R(x̂(r, tH )) − tL · C(x̂(r, tH ))
or
r · [R(x̂(r, tL )) − R(x̂(r, tH ))] ≥ tL · [C(x̂(r, tL )) − C(x̂(r, tH ))].
That implies — since 0 < r < 1 — that
R(x̂(r, tL )) − R(x̂(r, tH )) > tL · [C(x̂(r, tL )) − C(x̂(r, tH ))]
or
R(x̂(r, tL )) − tL · C(x̂(r, tL )) > R(x̂(r, tH )) − tL · C(x̂(r, tH ))
and hence (since tH > tL )
R(x̂(r, tL )) − tL · C(x̂(r, tL )) > R(x̂(r(tH ), tH )) − tH · C(x̂(r, tH ))
The term on the left of the inequality is W (r, tL ) and the term on the right is W (r, tH ). That
completes the proof.
Proof of Part (h)
Since x̂(rH , t) is a maximizer of rH · R(x) − t · C(x), we have
rH ·R(x̂(rH , t))−t·C(x̂(rH , t)) ≥ rH ·R(x̂(rL , t))−t·C(x̂(rL , t)) > rH ·R(x̂(rL , t))−t·C(x̂(rL , t)).
27
That implies (h).
Proof of Part (k)
Part (k) concerns the endogenous case. It says:
(1−r∗(tL ))·R(x̂(r∗(tL ), tL ) > (1−r∗(tH ))·R(x̂(r∗(tH ), tH )whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
When t drops from tH to tL , the Principal could continue to use the share r∗ (tH ). It suffices
to show that if he does so, the Principal’s gain cannot be less than it was at t = tH . Then, a
fortiori, it cannot be less when he uses r∗ (tL ), which is his best share when t = tL . Part ((b))
tells us that x̂(r∗ (tH ), tL ) ≥ x̂(r∗ (tH ), tH ). Since R is nondecreasing, we have
(1 − r∗ (tH )) · R(x̂(r∗ (tH ), tL ) ≥ (1 − r∗ (tH )) · R(x̂(r∗ (tH ), tH ).
That concludes the proof.
Proof of Part (l)
Part (l) concerns the endogenous case. It says:
W (r∗ (tL ), tL ) > W (r∗ (tH ), tH ) whenever tL , tH ∈ Γ̃ and 0 < tL < tH .
The effort x̂(r∗ (tL ), tL ) is a maximizer of r∗ (tL ) · R(x) − tL · C(x). Hence
r∗ (tL ) · R(x̂(r∗ (tL ), tL )) − tL · C(x̂(r∗ (tL ), tL )) ≥ r∗ (tL ) · R(x̂(r∗ (tH ), tH )) − tL · C(x̂(r∗ (tH ), tH ))
or
r∗ (tL ) · [R(x̂(r∗ (tL ), tL )) − R(x̂(r∗ (tH ), tH ))] ≥ tL · [C(x̂(r∗ (tL ), tL )) − C(x̂(r∗ (tH ), tH ))].
That implies — since 0 < r∗ (tL ) < 1 — that
R(x̂(r∗ (tL ), tL )) − R(x̂(r∗ (tH ), tH )) > tL · [C(x̂(r∗ (tL ), tL )) − C(x̂(r∗ (tH ), tH ))]
or
R(x̂(r∗ (tL ), tL )) − tL · C(x̂(r∗ (tL ), tL )) > R(x̂(r∗ (tH ), tH )) − tL · C(x̂(r∗ (tH ), tH ))
and hence (since tH > tL )
R(x̂(r∗ (tL ), tL )) − tL · C(x̂(r∗ (tL ), tL )) > R(x̂(r∗ (tH ), tH )) − tH · C(x̂(r∗ (tH ), tH ))
The term on the left of the inequality is W (r∗ (tL ), tL ) and the term on the right is W (r∗ (tH ), tH ).
That completes the proof.
2
Proof of Theorem 2
Part (a): For fixed t we have
d
R(x̂(r, t)) − t · C(x̂(r, t)) = R0 · x̂r − t · C 0 · x̂r = x̂r · (R0 − tC 0 ).
dr
28
But R0 − tC 0 > 0 since r · R0 − tC 0 = 0 and 0 < r < 1. Since we know from Part (a) of Theorem
d
1 that x̂r ≥ 0, we conclude that dr
W (r, t) ≥ 0.
Part (b): For fixed r we have
d
R(x̂(r, t)) − t · C(x̂(r, t)) = R0 · x̂t − t · C 0 · x̂t − C = x̂t · (R0 − tC 0 ) − C.
dt
Since we know from Part (b) of Theorem 1 that x̂t ≤ 0, we conclude (again using R0 − tC 0 > 0)
that dtd W (r, t) ≥ 0.
2
A proof of Theorem 3 which does not directly use the standard monotone-comparative-statics
proposition cited in the text.
Under the assumptions of the Theorem, the Agent-effort function x̂ is twice differentiable
with respect to both of its arguments at any (r, t) ∈ Γ. Consider the function x̂t and any interval
[r, 1), where 0 < r < 1. By the Mean Value Theorem, there exists r0 ∈ [r, 1) such that
x̂t (1, t) − x̂t (r̂, t) = x̂rt (r0 , t) · (1 − r).
Since r < 1, we conclude that
d
d
[x̂(1, t) − x̂(r, t)] > 0 at every (r, t) ∈ Γ ( [x̂(1, t) − x̂(r, t)] < 0 at every (r, t) ∈ Γ)
dt
dt
if and only if
x̂rt (r, t) > 0 at every (r, t) ∈ Γ (x̂rt (r, t) < 0 at every (r, t) ∈ Γ).
2
That is what Theorem 3 asserts.
Proof of Theorem 4
The function x̂ is an implicit function determined by the equation rR0 (x) − tC 0 (x) = 0 on
[0, J]. Since C 0 , R0 are twice differentiable, the function x̂ is also twice differentiable with respect
to r, t at all (r, t) in the convex set Γ, and so is the function W . Since x̂ and W are twice
differentiable, we have x̂rt = x̂tr andWrt = Wtr .
But Theorem 3 tells us that the statement (+) is equivalent to the statement
x̂rt (r, t) < 0 at every (r, t) ∈ Γ (x̂rt (r, t) > 0 at every (r, t) ∈ Γ).
By a Mean-Value-Theorem argument analogous to the one we used in proving Theorem 3, it is
also true that (++) is equivalent to the statement
Wrt (r, t) < 0 at every (r, t) ∈ Γ (Wrt (r, t) > 0 at every (r, t) ∈ Γ).
Hence the statement
(+) ⇐⇒ (++)
29
(which the Theorem asserts) is equivalent to the statement
x̂rt (r, t) · Wrt (r, t) > 0 whenever (r, t) ∈ Γ.
(1)
We now establish (1). We shall use the fact that the cross partials Wrt , Wtr are equal.
We know from earlier results that x̂r (r, t) ≥ 0. Since rR0 (x̂(r, t)) − tC 0 (x̂(r, t)) = 0 and
0 < r < 1, we have
Wr (r, t) = x̂r (r, t) · [R0 (x̂(r, t)) − tC 0 (x̂(r, t))] ≥ 0.
(2)
Since, by assumption, tC 0 (x̂(r, t)) = rR0 (x̂(r, t)), the equality in (2) can be rewritten
Wr (r, t) = (1 − r) · R0 (x̂(r, t)) · x̂r (r, t).
(3)
Now differentiate both sides of (3) with respect to t. We have:
(4)
Wrt (r, t) = (1 − r) · [R0 (x̂(r, t) · x̂rt (r, t) + x̂r (r, t) · R00 (x̂(r, t)) · x̂t (r, t)].
Next we differentiate W in the reverse order: first with respect to t and then with respect to
r. We obtain the following, using condensed notation when convenient:32
(5)
Wt (r, t) = R0 · x̂t − [t · C 0 · x̂t + C(x̂(r, t))] = x̂t · [R0 − tC 0 ] − C(x̂(r, t)).
Differentiating the final expression with respect to r, we obtain:
(6)
Wtr (r, t) = x̂t · [R00 · x̂r − t · C 00 · x̂r ] + [R0 − t · C 0 ] · x̂tr − C 0 · x̂r .
We now rewrite (4) and (6). We identify separate terms so that cancellations in the equality
Wrt = Wtr can be easily detected. For (4) we obtain
Wrt = R0 · x̂rt + x̂r · R00 · x̂t −r · R0 · x̂rt − r · x̂r · R00 · x̂r .
| {z } | {z }
1
2
For (6) we obtain
Wtr = x̂t · R00 · x̂r −t · C 00 · x̂r · x̂t + R0 · x̂tr −t · C 0 · x̂tr − C 0 · x̂r .
| {z }
| {z }
2
Deleting the terms
equality Wrt = Wtr as
1
and
1
2
and multiplying both sides by −1, we can now write the
r · R00 · x̂rt + r · x̂r · R00 · x̂t = t · C 00 · x̂r · x̂t + t · C 0 · x̂tr + C 0 · x̂r
Thus
R, R0 , R00 , C, C 0 , C 00 , x̂t , x̂r , x̂rt , Wr , Wt , Wrt
denote,
respectively,
, .... , C(x̂(r, t)) , .... , x̂t (r, t), x̂r (r, t), x̂rt (r, t), Wr (r, t), Wt (r, t), Wrt (r, t).
32
30
R(x̂(r, t)), R0 (x̂(r, t))
or
x̂rt · [rR0 − tC 0 ] +x̂r x̂t r · [R00 − t · C 00 ] = C 0 · x̂r .
| {z }
=0
So
C 0 · x̂r = x̂t · x̂r · (R00 − tC 00 ).
(7)
Now, using (7), we can rewrite (6) as
Wtr = x̂t · x̂r · [R00 − tC 00 ] + [R0 − t · C 0 ] · x̂tr − x̂t · x̂r · [R00 − tC 00 ] = [R0 − t · C 0 ] · x̂tr .
But R0 − t · C 0 > 0. So Wrt has the same sign as x̂rt . That establishes (1) and completes the
proof.
2
Proof of Theorem 5
Part (a)
We have
0
∂
∗
W (1, t) − W (r (t), t) = Wt (1, t) − Wr (r∗ (t), t) · r∗ (t) − Wt (r∗ (t), t).
∂t
(1)
Our exogenous Theorem 4 tells us that for all (r, t) — including the pair (r∗ (t), t) — we have
Wrt (r, t) · x̂rt (r, t) ≥ 0. That implies (since we assume that x̂rt (r, t)) < 0 at all (r, t) ∈ Γ) that
Wrt (r∗ (t), t) ≤ 0. Now consider (1) the interval [r∗ (t), 1], (2) the function Wt on the possible
shares, and (3) the derivative of that function, namely Wrt . By the Mean Value Theorem, there
exists r0 ∈ (r∗ (t), 1) such that
Wt (1, t) − Wt (r∗ (t), t) = Wrt (r0 , t) · (1 − r∗ (t)) ≤ 0
and hence
Wt (1, t) − Wt (r∗ (t), t) ≤ 0.
(2)
Since we know, from Part (g) of Theorem 1, that Wr (r, t) ≥ 0 at all (r, t), including (r∗ (t), t),
0
and since we assume r∗ (t) ≥ 0, we conclude, using (1) and (2), that
∂
∗
W (1, t) − W (r (t), t) ≤ 0.
(3)
∂t
We now turn to
(4)
∂
∂t
∗
x̂(1, t) − x̂(r (t), t) . We have
0
∂
∗
x̂(1, t) − x̂(r (t), t) = x̂t (1, t) − x̂t (r∗ (t), t) − x̂r (r∗ (t), t)) · r∗ (t).
∂t
31
Our assumption that x̂rt (r∗ (t), t) < 0 implies, using Theorem 3, that x̂t (1, t) − x̂t (r∗ (t), t) < 0.
0
Since we assume that r∗ (t) > 0, and
since we know from
Part (a) of Theorem 1, that x̂r (r, t) ≥ 0
at all (r, t), we conclude that
∂
∂t
x̂(1, t) − x̂(r∗ (t), t) ≤ 0. So we indeed have
∂
∂
∗
∗
x̂(1, t) − x̂(r (t), t) ·
W (1, t) − W (r (t), t) ≥ 0.
∂t
∂t
Part (b)
First consider again the three terms at the right of the equality (1). Since we now assume
x̂rt (r, t) ≥ 0, Theorem 4 tells us that we now have Wrt (r, t) ≥ 0. Hence — using the Mean-ValueTheorem argument again — we now have
Wt (1, t) − Wt (r∗ (t), t) ≥ 0.
0
Since we now assume r∗ (t) < 0, we conclude (using (1)) that
∂
∗
W (1, t) − W (r (t), t) ≥ 0.
∂t
Now consider the three terms on the right of the equality in (4). Our assumption that x̂rt (r∗ (t), t) ≥
0 now implies, using Theorem 3, that x̂t (1, t) − x̂t (r∗ (t), t) ≥ 0. Since we now assume that
0
r∗ (t) < 0, we now conclude, using (4), that x̂r (r, t) ≥ 0 and hence x̂t (1, t) − x̂t (r∗ (t), t) < 0. So
we again have
∂
∂
∗
∗
x̂(1, t) − x̂(r (t), t) ·
W (1, t) − W (r (t), t) ≥ 0.
∂t
∂t
2
Proof of Theorem 6
The theorem concerns the Principal’s gain (residual revenue), which we now denote H(r, t).
Thus
H(r, t) ≡ (1 − r) · R(x̂(r, t)).
We use the following fact about implicit functions: If we have f (ū, v̄) = 0, where f is twice
differentiable, then there exists a neighborhood of (ū, v̄) and a twice differentiable function g
such that for all (u, v) in the neighborhood we have f (g(u), v) = 0 and, if fv (g(ū), v̄) 6= 0, we
have
−fu (g(v), v)
g 0 (v) =
.
fv (g(v), v)
To apply this, let r play the role of u, let t play the role of v, let Hr play the role of f and
consider the first-order equation satisfied by r∗ (t):
Hr (r, t) = 0.
32
Under our assumptions, Hr is differentiable with respect to r and t at every (r, t) ∈ Γ and r∗ (t)
belongs to the open interval (0, 1) and is the unique solution to Hr (r, t) = 0. The role of the
twice differentiable function g is now played by r∗ .
Since r∗ (t) is the unique interior maximizer of H(r, t) it satisfies the second-order condition
Hrr (r, t) ≤ 0.
(+)
For every t ∈ Γ̃, we have
0
r∗ (t) =
−Hrr (r∗ (t), t)
if Hrt (r∗ (t), t) 6= 0.
Hrt (r∗ (t), t)
In view of the second-order condition (+), the numerator in this fraction is nonnegative. Hence,
0
if Hrt (r∗ (t), t) ≥ 0, then
(++)
0
0
r ∗ (t) > 0 if Hrr (r, t) < 0; r ∗ (t) = 0 if Hrr (r, t) = 0.
0
We now claim that Hrt (r∗ (t), t) > 0. To see this, note that
Hrt = (1 − r) · [R0 · x̂rt + x̂r · R00 · x̂t ] − R0 · x̂t .
By assumption we have x̂t < 0, x̂r > 0, R0 > 0, R00 < 0, and x̂rt ≥ 0. Hence −R0 · x̂t > 0 and
(1 − r) · [R0 · x̂rt + x̂r · R00 · x̂t ] ≥ 0.
0
We conclude that Hrt > 0 at every (r, t) ∈ Γ. That implies, in view of (++), that r ∗ (t) ≥ 0 at
every t ∈ Γ̃, as the Theorem asserts.
2
Proof of Theorem 7.
The derivative of effort gap with respect to t is
0
x̂t (1, t) − x̂t (r∗ (t), t) − x̂r (r∗ (t), t) · r∗ (t).
0
Consider the case where r∗ (t) ≥ 0 and xrt < 0 (Box 2). Since 1 > r∗ (t), it follows from Theorem
0
3 and x̂rt < 0 that the term x̂t (1, t) − x̂t (r∗ (t), t) is negative. The term −x̂r (r∗ (t), t) · r∗ (t) is
0
negative or zero, since r∗ (t) ≥ 0 and (by Part (a) of Theorem 1) x̂r ≥ 0. So, by Theorem 4, the
derivative of the Penalty (surplus gap) with respect to t is negative or zero.
0
Now consider the case where r∗ (t) < 0 and xrt ≥ 0 (Box 3). It follows from Theorem 3
0
and x̂rt ≥ 0 that the term x̂t (1, t) − x̂t (r∗ (t), t) is nonnegative. The term −x̂r (r∗ (t), t) · r∗ (t) is
0
nonnegative, since r∗ (t) < 0 and x̂r ≥ 0. So, by Theorem 4, the derivative of the Penalty (surplus
gap) with respect to t is nonnegative. That completes the proof.
2
Proof of Theorem 8.
Part (a)
We first establish the following general proposition:
33
Let h(x) = (1 − x) · g(x). If g is increasing and concave on (0, 1), then h is concave
on (0, 1)
Here is the proof:
Consider x1 > 0 and x2 < 1. Consider λ ∈ (0, 1) and let x0 denote λx1 + (1 − λ) · x2 .
We shall prove that
h(x0 ) = (1 − x0 ) · g(x0 ) ≥ λ · (1 − x1 ) · g(x1 ) + (1 − λ) · (1 − x2 ) · g(x2 ).
We have
0 = λ · [(x0 − x1 ) + (1 − λ) · (x0 − x2 )] · g(x2 )
≥ λ · (x0 − x1 ) · g(x1 ) + (1 − λ) · (x0 − x2 ) · g(x2 )
= λ · (x0 − 1 + 1 − x1 ) · g(x1 ) + (1 − λ) · (x0 − 1 + 1 − x2 ) · g(x2 ).
Moving parts of the last expression to the left of the inequality we obtain:
λ · (1 − x1 ) · g(x1 ) + (1 − λ) · (1 − x2 ) · g(x2 ) ≤
=
≤
=
λ · (1 − x0 ) · g(x1 ) + (1 − λ) · (1 − x0 ) · g(x2 )
(1 − x0 ) · [λ · g(x1 ) + (1 − λ) · g(x2 )]
(1 − x0 ) · g(x0 )
h(x0 ).
The final inequality holds because g is concave.
Now apply this general proposition to our case, where r plays the role of “x” and R(x̂(r, t)
plays the role of “g(x)”. That establishes Part (a).
Part (b)
R(x̂(r, t)) is concave in r if its second derivative is negative, i.e.,
R0 · x̂rr + x̂r · R00 < 0.
That is the case, as claimed, if R00 < 0 and x̂rr < 0. To check the claimed sufficient condition for
x̂rr < 0, start by writing the first-order condition satisfied by x̂(r, t):
r · R0 (x̂(r, t)) − t · C 0 (x̂(r, t)) = 0.
Differentiating with respect to r we obtain
x̂r =
R0
.
tC 00 − rR00
Since x̂(r, t) is an interior maximizer, the denominator is negative. Differentiating both sides
with respect to r we obtain:
2
R0 · R00
1
00
00
00
0
000
000
·
[R
·
(tC
−
rR
)
+
R
·
(rR
−
tC
)]
+
.
x̂rr =
tC 00 − rR00
(tC 00 − rR00 )2
34
Since tC 00 −rR00 < 0 and, by assumption, R00 < 0, we see, as claimed, that x̂rr < 0 if rR000 −tC 000 ≤
0.
2
Proof that the “exploding marginals” example lies in Box 3 of the Interior Examples table.
We shall show this by examining the construction of the example. Recall that Box 3 requires
0
that for every (r, t) ∈ Γ we have r∗ (t) < 0 and x̂rt (r, t) ≥ 0.
As long as R0 6= 0, the first-order condition r · R0 (x̂(r, t)) = t · C(x̂(r, t)) can be written
C 0 (x̂(r, t))
r
= .
0
R (x̂(r, t))
t
That defines an implicit function g which satisfies
x̂(r, t) = g
r
t
.
It will be convenient to let S denote the ratio rt . So
x̂(r, t) = g(S) and
We have:
C 0 (g(S))
= S.
R0 (g(S)
1
−r
x̂r = g 0 · ; x̂t = g 0 · 2 .
t
t
−r
1
x̂rt = x̂tr = g 00 · 3 − 2 · g 0 .
t
t
x̂rt > 0 ⇐⇒ g 00 ·
(a)
Since r∗ satisfies the first-order condition 0 =
[(1 − r) · R(x̂(r, t))], we have
R
R t
=
· .
R0 · x̂r
R0 g 0
1 − r∗ =
(b)
d
dr
r
+ g 0 < 0.
t
Using (b), we obtain:
0
r∗ (t) = −∆,
where
d
R t
∆ =
·
dt
R0 g 0
R
t
d
t
R
d
·
+
·
=
dt R0
g0
dt g 0
R0
1
t
R g 0 − g 00 · t ·
0 2
00
·
·
(R
·
=
)
−
R
·
R
·
x̂
+
t
(R0 )2 g 0
R0
(g 0 )2
35
−r
t2
Since x̂t ·
t
g0
r
t
= −S and
= S, we have:
∆ = −S ·
(c)
(R0 )2 − R00 · R
R
g 0 − g 00 · (−S)
+
+
.
(R0 )2
R0
(g 0 )2
To construct our example, we now let g(S) = ln ln S. Then
g0 =
1
1
; g 00 = −(ln S + 1) · 2
.
S · ln S
S · (ln S)2
[We can then verify that the condition in (a) is satisfied, and hence x̂rt > 0, as Box 3 requires].
Hence
−(ln S + 1)
ln S − (ln S + 1)
−1
1
0
00
=
+S·
=
.
(d)
g +g ·S =
2
2
2
S · ln S
S · (ln S)
S · (ln S)
S · (ln S)2
Using (c),(d), and the fact that (g 0 )2 =
∆ = −S ·
1
,
S 2 ·(ln S)2
we obtain
(R0 )2 − R00 · R
R
−S
−S· 0 =
· [(R0 )2 − R00 · R + RR0 ].
0
2
(R )
R
(R0 )2
0
Thus (recalling that r∗ = −∆) we have
0
r∗ < 0 ⇔ (R0 )2 − R00 · R0 + R · R0 < 0.
(e)
To continue our construction, we now suppose that
2
R = ekx , where k > 0.
We now claim that
(f)
2
(R0 )2 − R00 · R0 + R · R0 = e2kx · (2kx − 2k).
To show this we first note that
2
R0 = ekx · 2kx
and
2
2
R00 = ekx · 2k + 2kx · ekx · 2kx.
2
We then factor out the term e2kx in writing the following expressions.
2
2
2
R · R0 = ekx · ekx · 2kx = e2kx · 2kx.
2
(R0 )2 = e2kx · 4k 2 x2 .
36
2
2
2
R00 · R = 2k · ekx · [1 + 2kx2 ] · ekx = e2kx · [2k + 4k 2 x2 ].
So
2
2
(R0 )2 − R00 · R0 + R · R0 = e2kx · [4k 2 x2 − 2k − 4k 2 x2 + 2kx] = e2kx · (2kx − 2k)
and (f) is verified.
So, in view of (e), we have
0
2
r∗ < 0 as Box 3 requires, if e2kx · (2kx − 2k) < 0.
2
But if e2kx · (2kx − 2k) < 0, then x < 1. So our set of available efforts will be Σ = [0, 1).
Summarizing, we have
2
• R = ekx .
• x̂(r, t) = g(s) = ln ln s < 1 and hence S < ee .
We need to specify our set Γ of possible pairs (r, t). It will be the set
{(r, t) : 0 < r ≤ 1;
r
∈ (e, ee )}.
t
It remains to specify the function C. We seek a function C with the following property:
for every S we have C 0 (g(S)) = S · R0 (g(S)).
That can be rewritten as:
for every S we have C 0 (g(S)) = g g −1 (S) · R0 (g(S)).
R
Now let M denote G(s). Since C(M ) = C 0 (M ) dM , we have:
Z 0
−1
(gs)
C(M ) =
g g (M ) · R (M ) dM.
In our example
• g(S) = ln ln S.
M
• Hence, for any M , we have g −1 (M ) = ee .
2
2
• R(x) = ekx and R0 (x) = ekx · 2kx2 .
37
Thus, in our example, the equality (g) becomes:
Z h
i
eM
kM
C(M ) =
e · e · 2kM dM.
So for every x in our effort set σ = (0, 1] we have
Z x
ep kp
C(x) =
e · e · 2kp dp.
0
To summarize, our Box 3 example is as follows.
• Σ = (0, ].
• Γ = {(r, t) : 0 < r ≤ 1; rt ∈ (e, ee )}.
2
• R(x) = ekx , where k > 0.
Rx p
• C(x) = 0 ee · ekp · 2kp dp.
In the example provided in the text we have k = 1.
REFERENCES
Courtney, D. and T. Marschak (2009). “Inefficiency and complementarity in sharing games”,
Review of Economic Design 13, 7-43.
Garicano, L. and A. Prat (2011). “Organizational economics with cognitive costs”. Working
paper, London School of Economics.
Holmstrom, B. (1979). “Moral hazard and observability”, The Bell Journal of Economics, 10,
74-91.
Laffont, J-J. and D. Martimort (2002). The Theory of Incentives: the Principal -Agent Model,
Princeton University Press.
T. Marschak (2006). “Organization structure”, in T. Hendershott (ed.), Handbook of Economics
and Information Systems, Elsevier, 205-290.
T. Marschak, J.G. Shanthikumar, and J. Zhou (2017). “Does more information-gathering effort
raise or lower the average quantity produced?”. Journal of Mathematical Economics, Vol.
69, March 2017, 104-117.
N. Nissan, T. Roughgarden, E. Tardos, V. Vazirani (2007). Algorithmic Game Theory, Cambridge University Press.
T. Roughgarden (2005). Selfish Routing and the Price of Anarchy, MIT Press.
38
R.K. Sundaram (1996). A First Course in Optimization Theory, Cambridge University Press.
ACKNOWLEDGMENT:
We are grateful to Junjie Zhou for important preliminary work on our problem. In particular,
he supplied the monotone-comparative-statics proofs of parts (i) and (j) in Theorem 1.
39