Equation (24.8): the LDA log-likelihood ratio ½(x−μ₀)ᵀΣ⁻¹(x−μ₀) − ½(x−μ₁)ᵀΣ⁻¹(x−μ₁) equals ⟨w, x⟩ + b with w = Σ⁻¹(μ₁−μ₀), b = ½(μ₀ᵀΣ⁻¹μ₀ − μ₁ᵀΣ⁻¹μ₁)
ProvedUnderstandingML.lda_log_likelihood_ratiobayes-optimallinear-classifierlinear-discriminant-analysis
Equation (24.8). In the LDA setting the log-likelihood ratio becomes , which can be rewritten as where and . As a result, under the generative assumptions of LDA the Bayes optimal classifier is a linear classifier.
Formally: the identity for any symmetric matrix in the role of .
Preamble
import Definitions.Def_UnderstandingML_Generative open MeasureTheory ProbabilityTheory open scoped InnerProductSpace
Formal statement
namespace UnderstandingML
/-- **Equation (24.8)** (p. 348). Under the LDA assumptions the log-likelihood ratio
`½(x − μ₀)ᵀΣ⁻¹(x − μ₀) − ½(x − μ₁)ᵀΣ⁻¹(x − μ₁)` equals `⟨w, x⟩ + b` with `w = Σ⁻¹(μ₁ − μ₀)` and
`b = ½(μ₀ᵀΣ⁻¹μ₀ − μ₁ᵀΣ⁻¹μ₁)`, so the Bayes optimal classifier is linear. Stated for any symmetric
matrix `M` in the role of `Σ⁻¹`. -/
theorem lda_log_likelihood_ratio {d : ℕ} (M : Matrix (Fin d) (Fin d) ℝ) (hM : M.IsSymm)
(μ₀ μ₁ x : Fin d → ℝ) :
1 / 2 * dotProduct (x - μ₀) (M.mulVec (x - μ₀)) - 1 / 2 * dotProduct (x - μ₁) (M.mulVec (x - μ₁)) =
dotProduct (M.mulVec (μ₁ - μ₀)) x +
1 / 2 * (dotProduct μ₀ (M.mulVec μ₀) - dotProduct μ₁ (M.mulVec μ₁)) := by sorry
end UnderstandingML
Source
Shalev-Shwartz and Ben-David, Understanding Machine Learning: From Theory to Algorithms, Cambridge University Press 2014, doi:10.1017/CBO9781107298019, §24.3 p. 348, Equation (24.8)
Human review
Confirmed by the mission captain (proposal self-audit).
Confirmed by the moderator at approval.