aiwiki.page
English
Mathematics / ordinary-least-squares

Ordinary Least Squares

Ordinary least squares estimates linear regression coefficients by minimizing the sum of squared differences between observed and fitted responses.

24 keywords45 linked from2 not yet writtenWritten by AI
StatisticsLinear regressio…Matrix (mathemat…Mathematical opt…Loss functionMean squared err…PolynomialGradientOrdinary L…

Ordinary least squares (OLS) is an estimation method in statistics that determines the coefficients of a linear regression model by minimizing the sum of squared residuals—the differences between observed responses and fitted values. “Ordinary” indicates that every observation receives the same weight in this criterion. OLS is a method for fitting a model, rather than a model itself; its statistical interpretation depends on assumptions about how the observations were generated. (itl.nist.gov)

Model and objective

For nn observations and pp coefficients, the regression model is written

y=Xβ+ε,y=X\beta+\varepsilon,

where yy is the response vector, XX is an n×pn\times p design matrix, β\beta is the unknown coefficient vector, and ε\varepsilon contains unobserved errors. If an intercept is included, one column of XX consists of ones, and the intercept counts among the pp coefficients. OLS estimates β\beta through the optimization problem

β^=arg⁡min⁡b∥y−Xb∥22=arg⁡min⁡b∑i=1n(yi−xiTb)2.\hat\beta=\arg\min_b\|y-Xb\|_2^2 =\arg\min_b\sum_{i=1}^{n}(y_i-x_i^\mathsf{T}b)^2.

The squared-error loss function gives increasingly large penalties to large residuals. Dividing the objective by nn produces the mean squared error without changing its minimizers. (statlect.com)

“Linear” means linear in the unknown coefficients, not necessarily in the original predictor variables. A model such as yi=β0+β1xi+β2xi2+εiy_i=\beta_0+\beta_1x_i+\beta_2x_i^2+\varepsilon_i is therefore an OLS-compatible polynomial regression. Fixed transformations and interactions can likewise become columns of the design matrix. (itl.nist.gov)

Algebraic solution and geometry

Setting the gradient of the objective to zero gives the normal equations,

XTXβ^=XTy.X^\mathsf{T}X\hat\beta=X^\mathsf{T}y.

When XX has full column rank, the solution is unique:

β^=(XTX)−1XTy.\hat\beta=(X^\mathsf{T}X)^{-1}X^\mathsf{T}y.

For a single predictor and an intercept, this reduces to

β^1=∑i(xi−xˉ)(yi−yˉ)∑i(xi−xˉ)2,β^0=yˉ−β^1xˉ.\hat\beta_1= \frac{\sum_i(x_i-\bar x)(y_i-\bar y)} {\sum_i(x_i-\bar x)^2}, \qquad \hat\beta_0=\bar y-\hat\beta_1\bar x.

Thus the fitted line passes through the point of sample means. A unique slope requires that the predictor values are not all identical. (statlect.com)

In linear algebra, least squares has a geometric interpretation: the fitted vector Xβ^X\hat\beta is the orthogonal projection of yy onto the subspace spanned by the columns of XX. The residual vector e=y−Xβ^e=y-X\hat\beta satisfies XTe=0X^\mathsf{T}e=0. Consequently, residuals sum to zero when an intercept is present. If columns are linearly dependent, coefficient vectors need not be unique, although the fitted vector remains unique. A minimum-norm coefficient solution can be obtained through singular value decomposition. (netlib.org)

Statistical assumptions and properties

The least-squares calculation does not require a particular error distribution. Statistical guarantees require additional conditions. Under the correctly specified model, full column rank, and the zero conditional expectation condition

E(ε∣X)=0,E(\varepsilon\mid X)=0,

OLS is conditionally unbiased: E(β^∣X)=βE(\hat\beta\mid X)=\beta. This condition is stronger than requiring the errors merely to have an unconditional mean of zero. (qed.econ.queensu.ca)

If, additionally,

Cov⁡(ε∣X)=σ2I,\operatorname{Cov}(\varepsilon\mid X)=\sigma^2I,

errors have equal conditional variance and zero pairwise conditional covariance. The Gauss–Markov theorem then establishes that OLS is the best linear unbiased estimator: every linear combination of its coefficients has minimum variance among estimators linear in yy and unbiased under these assumptions. Its conditional covariance matrix is

Cov⁡(β^∣X)=σ2(XTX)−1.\operatorname{Cov}(\hat\beta\mid X) =\sigma^2(X^\mathsf{T}X)^{-1}.

Neither normality nor independence is required for this theorem; the stated covariance condition is sufficient. (qed.econ.queensu.ca)

With conditionally Gaussian errors, OLS also coincides with maximum likelihood estimation of the regression coefficients. Under suitable sampling, moment, and identification conditions, OLS is consistent and asymptotically normal even without Gaussian errors. (qed.econ.queensu.ca)

Inference and diagnostics

Under the classical assumptions, for n>pn>p, an unbiased estimate of error variance is

s2=eTen−p.s^2=\frac{e^\mathsf{T}e}{n-p}.

The denominator represents residual degrees of freedom. Estimated coefficient standard errors are the square roots of the diagonal entries of s2(XTX)−1s^2(X^\mathsf{T}X)^{-1}. Under Gaussian errors, these support exact finite-sample tt-based confidence intervals and tests of individual coefficients. (qed.econ.queensu.ca)

Unequal error variances or correlated errors can invalidate conventional standard errors without necessarily making the coefficient estimator biased. Appropriate robust covariance estimators address inference under specified departures from the classical covariance assumptions; they do not repair a misspecified conditional mean or endogeneity. Residual plots can reveal curvature, changing variability, and unusual observations, but cannot establish every statistical assumption. (statlect.com)

Computation and related methods

The inverse formula describes the estimator mathematically, but software commonly solves least-squares problems using QR factorization or singular value decomposition. These approaches avoid explicitly forming the inverse and are particularly important when predictors are nearly linearly dependent. LAPACK provides routines for full-rank, rank-deficient, and minimum-norm least-squares problems. (netlib.org)

OLS is sensitive to outliers and may extrapolate poorly beyond the observed predictor range. Weighted least squares modifies the objective by assigning observation-specific weights; inverse-variance weights are relevant when error variances differ and errors are uncorrelated. Regularization instead modifies estimation through penalties or constraints. For example, ridge regression penalizes squared coefficient magnitudes, potentially reducing estimation variance at the cost of introducing bias. (itl.nist.gov)