Markov Decision Processes XX: Perpetual American Options and Credit GrantingTextbook
Motivation
An American option may be exercised at any moment up to maturity, so pricing one is not an integration problem but a stopping problem: the holder must decide, at each date and in each state of the market, whether the payoff available now beats the option value of waiting. A perpetual American put pushes this to its limit — there is no maturity at all, so the horizon is unbounded and the problem has no terminal condition to induct backwards from. What replaces the terminal condition is a fixed point characterization, and the classical answer, going back to McKean (1965) and Merton (1973) in continuous time and to Cox, Ross and Rubinstein (1979) in the binomial model, is that the price is the smallest superharmonic majorant of the payoff.
Bäuerle and Rieder's Chapter 11 (Markov Decision Processes with Applications to Finance, Springer, 2011) derives this from their own general unbounded-horizon stopping theory rather than from stochastic analysis, and in the same chapter applies the bounded-horizon version to a problem from banking rather than trading: when should a bank cancel a credit line? The two halves share one mathematical shape — a stopping problem whose optimal policy turns out to be of threshold type — and this mission formalizes both, with the perpetual put as the goal.
Setting
The binomial model (§11.1). A stock moves from price to with risk-neutral probability and to with , where , and the discount factor is . The defining relation of the risk-neutral measure,
is carried as a hypothesis of the model: it is what makes the discounted stock price a martingale, and the proofs use it directly.
The American put with strike pays when exercised. With periods to maturity its price satisfies the recursion and
the maximum being "exercise now" against "hold". Proposition 11.1.2 describes the price at time of an option maturing at : it is continuous in , decreasing in , and — the part that carries the argument — is increasing, even though itself decreases in . That single reformulation, obtained by adding to both sides of the recursion and using the risk-neutral relation, is what yields the threshold structure: there are exercise boundaries with optimal. Exercise when the stock falls far enough, and the boundary rises as maturity approaches.
The perpetual put (Theorem 11.1.3, the goal). With no expiration date the price at time zero is a supremum over all stopping times, included:
with the stopping reward set to zero on . The theorem says four things: is the limit of the finite-maturity prices ; solves and satisfies ; is the smallest superharmonic function majorizing ; and, if the value of the exercise-region policy dominates , then equals that value and , the hitting time of , is optimal.
The conditional in part d) is the book's own and is not decoration: without it the exercise region need not deliver an optimal stopping time, and the unconditional version is a different, false statement. The boundedness in b) is likewise a genuine claim rather than a side remark — the fixed point equation alone admits other solutions, and it is boundedness together with minimality that pins down among them.
Credit granting (§11.2). A bank holds a credit contract of maximal duration . Each period it observes a rating class evolving as a Markov process , and chooses to extend — earning — or to cancel, ending the contract. The value iteration is . Under two structural assumptions — increasing, and stochastically monotone, so that a better-rated borrower stays better-rated — Theorem 11.2.1 gives the same shape of answer as the option: cancel exactly when the rating falls below a threshold , and those thresholds rise as the remaining duration shortens, since a marginal borrower is no longer worth keeping when there is little time left to recover.
Theorem 11.2.2 repeats this when the borrower is not rated at all. The bank has only a prior on the repayment probability and one signal per period; the state records positive signals out of , and the expected repayment probability is the posterior mean . Monotonicity here is for the order of p. 342 — more positive signals and fewer negative ones — and not the coordinatewise order, under which the claim would be false: an extra signal that is negative makes the state worse.
What is being asked
Formalize Theorem 11.1.3 in full, all four parts: the limit identification, the fixed point equation with its bounds, the minimality among superharmonic majorants, and the conditional optimality of the exercise-region stopping time. The three milestones are Proposition 11.1.2, the finite-horizon put whose threshold structure the perpetual case specializes, and the two credit granting theorems, which run the same bounded-horizon argument on a different model.
The stopping-time apparatus is built rather than assumed: the stock path, the law pinned by its finite-dimensional distributions, stopping times valued in , and the reward vanishing at . Part a) of the goal is the identification of the supremum with , so carrying as an abstract function would make the theorem vacuous.