A covariance matrix is a square matrix that records the variances of several random variables and the covariances between every pair. In statistics, it summarizes how variables fluctuate individually and together. Its diagonal entries measure individual variability, while its off-diagonal entries measure whether deviations from the variables’ means tend to have the same or opposite signs. It is also called a variance–covariance matrix or dispersion matrix. (itl.nist.gov)
Definition and interpretation
Let be a real-valued random vector whose components have finite second moments, and let be its vector of expected values. Its covariance matrix is
where denotes matrix transposition. Equivalently,
Consequently, . (statlect.com)
A positive off-diagonal entry indicates that two variables tend to deviate from their means in the same direction; a negative entry indicates opposite-direction deviations. Covariance depends on measurement units: its units are the product of the units of the two variables. Therefore, covariance magnitudes cannot generally be compared across differently scaled variable pairs. (online.stat.psu.edu)
For illustration,
describes two variables with standard deviations and . Their Pearson correlation coefficient is . This example illustrates the distinction between covariance, which retains scale, and correlation, which standardizes it. (online.stat.psu.edu)
Algebraic properties
A real covariance matrix is symmetric and positive semidefinite. For every deterministic vector ,
Thus, the matrix determines the variance of every linear combination of the variables. It is positive definite precisely when no nonzero linear combination has zero variance. Otherwise, some nonzero linear combination of the centered variables equals zero almost surely. (statlect.com)
All its eigenvalues are nonnegative. The spectral theorem gives a decomposition
where is an orthogonal matrix and contains the eigenvalues. Zero eigenvalues indicate directions with no variability. Such a matrix is singular and has no ordinary matrix inverse. (statlect.com)
The trace is the sum of the component variances and also the sum of the eigenvalues. In principal-component analysis, this quantity represents total variance in the chosen coordinate system; its numerical value depends on variable scaling. (itl.nist.gov)
Transformations and correlation
For a deterministic matrix and constant vector , the affine transformation satisfies
Adding constants therefore leaves covariance unchanged, while rescaling or mixing variables changes it predictably. This identity provides an exact rule for propagating covariance through linear transformations. (statlect.com)
If all component variances are positive, define
The corresponding correlation matrix is
Its diagonal entries equal one, and its off-diagonal entries are dimensionless Pearson correlations. This is the covariance matrix of variables standardized by subtracting their means and dividing by their standard deviations. The distinction matters in feature scaling, because covariance-based analyses retain the original measurement scales. (online.stat.psu.edu)
Estimation from observations
Given independent, identically distributed observations , let denote their sample mean. The conventional unbiased estimator is
The denominator implements Bessel’s correction for estimating the mean from the same observations. If is the data matrix with each column centered, the equivalent expression is . (itl.nist.gov)
Under a multivariate Gaussian model with unknown mean, maximum likelihood estimation instead uses denominator . Empirical covariance estimates can be unstable when the number of variables is large relative to the number of observations, particularly when an inverse is required. (scikit-learn.org)
One form of regularization is shrinkage:
where is the identity matrix. This blends the empirical estimate with a spherical target, trading some estimation bias for improved stability. Robust covariance estimators address a different difficulty: sensitivity to outlying observations. (scikit-learn.org)
Statistical uses and limitations
In principal component analysis, covariance eigenvectors identify mutually orthogonal directions, and their eigenvalues give the variances along those directions. Retaining directions with the largest eigenvalues provides dimensionality reduction while preserving as much variance as possible for the retained dimension. Using a correlation matrix instead amounts to performing the analysis on standardized variables. (itl.nist.gov)
For a multivariate normal distribution, the mean vector and covariance matrix completely specify the distribution. With nonsingular covariance, its density has ellipsoidal contours determined by the quadratic form
This expression also accounts for different scales and correlations between coordinates. (itl.nist.gov)
Outside special distributional families, covariance does not determine the full joint distribution. Zero covariance means absence of linear association, not necessarily statistical independence. For jointly Gaussian variables, however, zero cross-covariances do imply independence. Consequently, a diagonal covariance matrix permits an independence interpretation under joint Gaussianity, but not for arbitrary data. (data140.org)