Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

§2 — V(f, π) = L(f)V(π) and V(f₁, ⋯, f_N, π) = L(f₁)⋯L(f_N)V(π)

Proved
BlackwellDiscreteDP.Stationary.composition_rule

by mikedeng1 · Oct 4, 2026 · Mathlib 0df444a (Lean v4.33.1)

dynamic-programmingmarkov-decision-processp2o-batch-pfp2ap2o-gran-per-chapterp2o-plan-paperp2o-v1

Fix a discount factor 0≤β<10\le\beta<10≤β<1 and recall L(f)w=r(f)+βQ(f)wL(f)w=r(f)+\beta Q(f)wL(f)w=r(f)+βQ(f)w. For every decision rule fff and policy π\piπ,

Vβ(f,π)=L(f) Vβ(π),V_\beta(f,\pi)=L(f)\,V_\beta(\pi),Vβ​(f,π)=L(f)Vβ​(π),

and more generally, for decision rules f1,…,fNf_1,\dots,f_Nf1​,…,fN​,

Vβ(f1,…,fN,π)=L(f1)⋯L(fN) Vβ(π).V_\beta(f_1,\dots,f_N,\pi)=L(f_1)\cdots L(f_N)\,V_\beta(\pi).Vβ​(f1​,…,fN​,π)=L(f1​)⋯L(fN​)Vβ​(π).

This is the recursion that underlies Theorems 1 and 2: putting a decision rule in front of a policy acts on returns by the affine monotone map L(f)L(f)L(f).

Formalization Note. L(f1)⋯L(fN)Vβ(π)L(f_1)\cdots L(f_N)V_\beta(\pi)L(f1​)⋯L(fN​)Vβ​(π) is the right fold of the list [f1,…,fN][f_1,\dots,f_N][f1​,…,fN​]; for the empty list both sides equal Vβ(π)V_\beta(\pi)Vβ​(π).

Preamble
import Mathlib
import Definitions.Def_BlackwellDiscreteDP_Stationary_Model
Formal statement
namespace BlackwellDiscreteDP.Stationary

/-- §2, p. 720 (unnumbered; Blackwell, *Discrete Dynamic Programming*, Ann. Math. Statist. 33(2):719–726 (1962),
DOI 10.1214/aoms/1177704593): the composition rule
`V(f, π) = L(f)V(π)` and `V(f₁, ⋯, f_N, π) = L(f₁) ⋯ L(f_N)V(π)`, for `0 ≤ β < 1`.

**Formalization Note.** `L(f₁) ⋯ L(f_N)` applied to `V(π)` is the right fold of the list
`[f₁, …, f_N]`, i.e. `L(f₁)(L(f₂)(⋯ L(f_N)(V(π))))`; the empty list gives `V(π)` itself. -/
theorem composition_rule {St Act : Type} [Fintype St] [DecidableEq St] [Nonempty St] [Fintype Act] [Nonempty Act]
    (M : Model St Act)
    (β : ℝ) (hβ0 : 0 ≤ β) (hβ1 : β < 1) :
    (∀ (f : St → Act) (π : Policy St Act),
        M.V β (Policy.cons f π) = M.L β f (M.V β π)) ∧
    (∀ (fs : List (St → Act)) (π : Policy St Act),
        M.V β (Policy.prepend fs π) = fs.foldr (fun f w => M.L β f w) (M.V β π)) := by sorry

end BlackwellDiscreteDP.Stationary
Source
Blackwell, Discrete Dynamic Programming, Ann. Math. Statist. 33(2):719–726 (1962), DOI 10.1214/aoms/1177704593, p. 720, §2 (unnumbered display text)
Human review
  • Endorsed by Shuze Chen · Oct 5, 2026

    Confirmed by the moderator at approval.

  • Endorsed by mikedeng1 · Oct 5, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me