Ordinary Least Squares and Regularization
For the model, residuals, and step-by-step gradient updates, start with linear regression. Here the question is how to solve the least-squares problem directly, when its coefficients are identifiable, and which additional assumptions support uncertainty estimates.
For design matrix , targets , and coefficient vector , ordinary least squares (OLS) solves
The pseudoinverse gives a least-squares solution, . When columns of are nearly linearly dependent, small changes in the data can produce large coefficient changes. Prediction may remain acceptable while individual coefficients become unstable.
From Residuals to a Solution
Here has rows (observations) and columns (coefficients), , and . Include a column of ones if an intercept is fitted. For a line , the residual is : a vertical difference in target units, not the shortest distance to the line. That perpendicular distance is in unscaled Cartesian coordinates; minimizing it would solve a different problem that also permits movement in .
For any point, keep x fixed and move vertically to the line: the point gives the observed y, and the line gives the predicted ŷ. Each blue segment shows the magnitude of a residual, rather than the perpendicular distance to the line. OLS chooses the line that minimizes the sum of these squared vertical differences. This D2L illustration uses its own data; the three-observation calculation below is a separate example.
Using leaves the OLS minimizers unchanged and gives
The Hessian is positive semidefinite, so a zero gradient is a global minimum. Setting it to zero gives the normal equations:
If (thus ), the solution is unique and equals . Otherwise coefficients are not unique: adding any vector in the null space of leaves the fitted values unchanged. The Moore–Penrose pseudoinverse selects the minimum-Euclidean-norm coefficient vector; the training fitted values are still unique. For example, duplicate columns identify only the sum of their coefficients. Predictions outside that column relationship need not agree.
Use a QR- or SVD-based least-squares solver rather than explicitly forming an inverse; forming squares the spectral condition number for full-column-rank . The scikit-learn linear-model guide describes OLS and its SVD solution.
Three Observations by Hand
Take . The intercept column and column are independent. With and ,
Predictions are , residuals are , and the squared-error sum is . Both and check the normal equations. The training-mean baseline predicts and has squared-error sum . This is a training comparison, not evidence of generalization.
from math import isclose
from statistics import mean
x, y = [0, 1, 2], [1, 2, 2]
xbar, ybar = mean(x), mean(y)
sxx = sum((v - xbar) ** 2 for v in x)
if sxx == 0:
raise ValueError("A constant x cannot identify slope and intercept separately")
a = sum((u - xbar) * (v - ybar) for u, v in zip(x, y)) / sxx
b = ybar - a * xbar
r = [v - (b + a * u) for u, v in zip(x, y)]
assert isclose(a, 0.5) and isclose(b, 7 / 6)
assert isclose(sum(v * v for v in r), 1 / 6)
assert isclose(sum(r), 0, abs_tol=1e-12)
assert isclose(sum(u * v for u, v in zip(x, r)), 0, abs_tol=1e-12)
Fitting Is Not Statistical Inference
No Gaussian assumption is needed to compute the minimizer. For statistical use, the statsmodels regression documentation makes the error covariance model explicit; choosing an uncertainty formula requires that extra model. Under a correctly specified conditional mean with and full column rank, OLS is conditionally unbiased. If also , then
For , residual sum of squares divided by estimates under these assumptions. Conditional Gaussian errors additionally justify the usual exact finite-sample t/F inference. Heteroskedastic or correlated errors require appropriate uncertainty estimators; robust standard errors do not fix a wrong conditional mean or confounding. An interval for a mean response is not a prediction interval for a new noisy observation.
Large residuals can dominate squared error, and high-leverage observations (unusual feature values) can strongly move the fitted line. Inspect residuals against fitted values and time, and compare held-out errors with a baseline before interpreting coefficients.
Penalized Objectives
Below, and . For these penalized formulas, use training-centered and let contain slopes only; recover the unpenalized intercept as . Scale features within training folds because coefficient penalties depend on units. Normalization is part of the definition: scikit-learn Ridge uses , so the Ridge formula here corresponds to . Its Lasso and Elastic Net conventions match , with l1_ratio equal to . The Ridge term at in the Elastic Net formula has half the strength of the Ridge formula below at the same .
Ridge
Ridge shrinks correlated or weakly identified coefficients and usually keeps all of them nonzero.
Lasso
The penalty can set coefficients exactly to zero. That sparsity is a property of the fitted objective, not proof that the selected variables are uniquely important or causal. With strongly correlated features, the selected member may be unstable.
Elastic Net
Elastic Net combines sparsity with Ridge-style stabilization and is often useful when predictors occur in correlated groups.
Working Rules
- Fit scaling and basis transformations on training data only; apply the recorded transformation to validation and test data.
- Do not penalize the intercept unless the formulation explicitly intends it.
- Select and using validation or cross-validation inside the training process.
- Compare predictive error, coefficient stability, and operational simplicity—not training loss alone.
- Use robust or quantile objectives when squared error does not represent the desired target or error cost.
- Separate predictive modeling from inferential claims; uncertainty estimates require explicit assumptions about the data-generating process.
Polynomial and interaction features can make the input relationship nonlinear while the fitted coefficients remain linear. The resulting model still inherits extrapolation and overfitting risks from the chosen basis.
See the scikit-learn linear-model guide for current solvers and estimator APIs.