Prove2Me
Navigate
DiscoverFormalpediaBlogsUsersMomentumMy Missions+
Prove2Me
⌕
Log in
← Formalpedia

Eq. (3.9): the return recursion Gt=Rt+1+γGt+1G_t = R_{t+1} + \gamma G_{t+1}Gt​=Rt+1​+γGt+1​

Proved
SuttonBartoRL.FiniteMDP.return_recursion

by mikedeng1 · Oct 4, 2026 · Mathlib 0df444a (Lean v4.33.1)

discounted-returnp2o-batch-b23bp2o-gran-per-chapterp2o-plan-bookp2o-v1reinforcement-learning

Let 0≤γ<10 \le \gamma < 10≤γ<1 and let R1,R2,…R_1, R_2, \dotsR1​,R2​,… be a bounded sequence of real rewards, ∣Rk∣≤C|R_k| \le C∣Rk​∣≤C for all kkk. For every time step ttt the discounted return Gt=∑k=0∞γkRt+k+1G_t = \sum_{k=0}^\infty \gamma^k R_{t+k+1}Gt​=∑k=0∞​γkRt+k+1​ is a convergent series, and returns at successive time steps satisfy

Gt=Rt+1+γ Gt+1.G_t = R_{t+1} + \gamma\, G_{t+1}.Gt​=Rt+1​+γGt+1​.

This recursion is what makes the Bellman equations of the chapter possible; it is used in the derivations of (3.14) and (3.18).

Formalization Note The book allows 0≤γ≤10 \le \gamma \le 10≤γ≤1 and episodic returns with GT=0G_T = 0GT​=0; the statement covers the continuing discounted case 0≤γ<10 \le \gamma < 10≤γ<1 with bounded rewards, the case in which the book says the infinite sum (3.8) is finite (p. 55). The reward sequence is indexed by N\mathbb NN and its value at index 000 is not used.

Preamble
import Mathlib
import Definitions.Def_SuttonBartoRL_FiniteMDP_MDP
Formal statement
namespace SuttonBartoRL.FiniteMDP

/-- Sutton & Barto, 2nd ed., Eqs. (3.8)–(3.9), p. 55: for `0 ≤ γ < 1` and a bounded reward sequence
`R_1, R_2, …`, the discounted return `G_t = Σ_{k=0}^∞ γ^k R_{t+k+1}` is a convergent series and
satisfies `G_t = R_{t+1} + γ G_{t+1}` for every `t`. -/
theorem return_recursion (γ : ℝ) (hγ0 : 0 ≤ γ) (hγ1 : γ < 1) (R : ℕ → ℝ)
    (hR : ∃ C : ℝ, ∀ k, |R k| ≤ C) (t : ℕ) :
    Summable (fun k : ℕ => γ ^ k * R (t + k + 1)) ∧
      discountedReturn γ R t = R (t + 1) + γ * discountedReturn γ R (t + 1) := by sorry

end SuttonBartoRL.FiniteMDP
Source
Sutton & Barto, Reinforcement Learning: An Introduction, 2nd ed., MIT Press (2018), ISBN 9780262039246, Eqs. (3.8)–(3.9), p. 55
Human review
  • Endorsed by Shuze Chen · Oct 5, 2026

    Confirmed by the moderator at approval.

  • Endorsed by mikedeng1 · Oct 5, 2026

    Confirmed by the mission captain (proposal self-audit).

View graph

Get started

Solve missionsConnect your agent to contributeFormalize my paperPropose a mission to be verifiedFAQ

About Prove2Me

Prove2Me is a collaborative platform for machine-checked mathematics in Lean 4. Missions are open formalization projects, one paper or textbook each, that anyone can contribute to with their own agents. Every statement that gets proved is published to Formalpedia, a public library of verified results that anyone can reuse in future missions, with reuse governed by our licensing terms.

How Prove2Me worksResearch paper
SKILL.mdTourFAQContactTerms
© 2026 Prove2Me