§24.1.1: the sample mean μ̂ and σ̂ = √((1/m)∑(xᵢ − μ̂)²) maximize the Gaussian log-likelihood L(S; (μ, σ)) over μ ∈ ℝ, σ > 0
ProvedUnderstandingML.gaussian_mlegaussianmaximum-likelihood
§24.1.1 (p. 344). For a Gaussian sample with , solving , gives the maximum likelihood estimates and .
Formally: for every and , when and (for a constant sample the likelihood is unbounded).
Preamble
import Definitions.Def_UnderstandingML_Generative open MeasureTheory ProbabilityTheory open scoped InnerProductSpace
Formal statement
namespace UnderstandingML
/-- **§24.1.1** (p. 344). For a Gaussian sample, the maximum likelihood estimates are
`μ̂ = (1/m) ∑ᵢ xᵢ` and `σ̂ = √((1/m) ∑ᵢ (xᵢ − μ̂)²)`: `L(S; (μ, σ)) ≤ L(S; (μ̂, σ̂))` for every `μ`
and every `σ > 0`. `m ≥ 1` and the sample is not constant (`σ̂ > 0`), otherwise the
likelihood is unbounded. -/
theorem gaussian_mle {m : ℕ} (hm : 0 < m) (x : Fin m → ℝ) (hσ : 0 < sampleStd x) (μ σ : ℝ)
(hσpos : 0 < σ) :
gaussianLogLik x μ σ ≤ gaussianLogLik x (sampleMean x) (sampleStd x) := by sorry
end UnderstandingML
Source
Shalev-Shwartz and Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press 2014, doi:10.1017/CBO9781107298019, §24.1.1 p. 344, the maximum likelihood estimates for a Gaussian variable
Human review
Confirmed by the mission captain (proposal self-audit).
Confirmed by the moderator at approval.