Linear and Quadratic Discriminant Analysis
Linear discriminant analysis (LDA) and quadratic discriminant analysis (QDA) model each class with a Gaussian distribution and combine the class-conditional density with a class prior .
For class ,
Ignoring terms shared by all classes, the discriminant score is
The predicted class has the largest score.
From Training Data to a Decision
Bayes' rule gives the posterior:
Estimate each mean from its training class, and estimate priors as unless justified target-population priors are supplied. One explicit maximum-likelihood convention estimates by dividing the within-class sum of outer products by ; LDA pools those sums across classes and divides by . Unbiased covariance conventions instead use or ; state the convention, since implementations need not match this calculation.
For a one-dimensional example, class 0 has training values and class 1 has . The fitted means are 0 and 2, the shared ML variance is 1, and both priors are . Hence
The boundary is ; at the posterior for class 1 is . With the same means and variance but priors , , the boundary moves to . A prior is not merely a label attached after fitting.
Choosing the largest posterior minimizes expected zero-one loss under the fitted model. Unequal error costs require minimizing posterior expected cost instead. Neither a Gaussian fit nor this decision rule guarantees that the model matches the deployment distribution.
LDA versus QDA
With a shared covariance, the quadratic terms in cancel between class scores, producing a linear boundary. Class-specific covariances retain those terms and produce a quadratic boundary. Class-specific diagonal covariance corresponds to a Gaussian naive Bayes-style conditional independence assumption.
Open full-size imageCompare columns within each row, then move downward as the covariance structure changes. Ellipses show estimated class spread. In the bottom row, LDA retains a straight boundary while QDA bends around classes with different covariances. Greater flexibility alone does not guarantee better predictions on new data.
Covariance Is the Hard Part
- High-dimensional or small-sample settings can make empirical covariance matrices noisy or singular.
- Shrinkage trades some bias for more stable covariance estimates and must be selected or estimated deliberately.
- Unregularized full-covariance LDA/QDA is invariant to a common invertible affine feature transformation in exact arithmetic. Scaling still affects numerical conditioning and some regularization schemes; class imbalance affects estimated priors.
- Closed-form parameter estimates do not mean the modeling choices require no validation.
The displayed density assumes positive-definite covariances. With features, an empirical class covariance has rank at most ; QDA therefore needs at least observations per class even to make full rank possible, and that count is not a guarantee of stability. A pooled LDA covariance has rank at most . Reduce redundant dimensions or use validated covariance regularization when these matrices are singular.
Supervised Projection
LDA can also project data onto directions that separate class means relative to within-class variation. This supervised dimensionality-reduction view is related to, but distinct from, using LDA as a classifier. With classes, at most discriminant directions carry between-class separation.
Gaussian assumptions can be useful approximations, but predicted probabilities still require calibration checks under the target distribution. Compare LDA and QDA with logistic regression and simple baselines rather than selecting by boundary shape alone.
See the scikit-learn LDA/QDA guide for current covariance estimators and implementation details.