Corollary 11.8: under independent priors, GREEDY suffers Bayesian regret
ProvedIntroBandits.greedy_linear_bayesian_regretbayesian-greedybayesian-regretincentivized-exploration
Corollary 11.8. Consider independent priors such that . Pick any such that . Then GREEDY suffers Bayesian regret
Formally: a prior supported on the finite under which the coordinates are independent (iIndepFun), , with , any GREEDY policy and any horizon; then , with the Bayesian regret (3.1) of Chapter 3 (the pseudo-regret in expectation over the run and the prior). The first hypothesis only guarantees that such an exists; it is kept as printed.
Preamble
import Definitions.Def_IntroBandits_Agents open MeasureTheory ProbabilityTheory BanditAlgorithm
Formal statement
namespace IntroBandits
theorem greedy_linear_bayesian_regret (P : Measure (Fin 2 → ℝ)) [IsProbabilityMeasure P]
(F : Finset (Fin 2 → ℝ)) (hF : P (↑F)ᶜ = 0)
(hunit : ∀ μ ∈ F, ∀ a, μ a ∈ Set.Icc (0 : ℝ) 1) (fam : RewardFamily)
(hindep : iIndepFun (fun (a : Fin 2) (μ : Fin 2 → ℝ) ↦ μ a) P)
(h1 : (P {μ | μ 0 = 1}).toReal < (priorMean P F 0 - priorMean P F 1) / 2)
{α : ℝ} (hα : 0 < α)
(hα2 : (P {μ | 1 - 2 * α ≤ μ 0}).toReal ≤ (priorMean P F 0 - priorMean P F 1) / 2)
{π : BanditPolicy 2} (hπ : IsGreedy P F fam π) (T : ℕ) :
T * (α / 2 * (priorMean P F 0 - priorMean P F 1) * (P {μ | 1 - α < μ 1}).toReal) ≤
bayesianRegret P F fam π T := by sorry
end IntroBandits
Source
Slivkins, Introduction to Multi-Armed Bandits, arXiv:1904.07272 (FnT ML 12, 2019), §11.2 p. 148, Corollary 11.8 with its proof (also Exercise 11.1)
Human review
Confirmed by the mission captain (proposal self-audit).
Confirmed by the moderator at approval.