Covariance is a numerical measure of joint variability in probability theory and statistics. It describes how the deviations of two random variables from their respective means vary together. Positive covariance indicates that deviations in the same direction predominate, while negative covariance indicates that deviations in opposite directions predominate. Unlike a standardized correlation coefficient, its magnitude depends on the variables’ measurement scales. (online.stat.psu.edu)
Definition and interpretation
For real-valued random variables and , with means and , covariance is defined by
Here, denotes expected value. Expanding the product gives the equivalent identity
The expectation is taken over the variables’ joint probability distribution, rather than their separate marginal distributions alone. For discrete variables it is a probability-weighted sum; for continuous variables with a joint density, it is a double integral. (online.stat.psu.edu)
Finite second moments, and , guarantee that covariance is finite. Each centered product is positive when both variables are above their means or both are below them, and negative when one is above its mean and the other below. Covariance averages these contributions, including their magnitudes. Zero covariance therefore means that the positive and negative contributions balance, not that the variables have no relationship. (ocw.mit.edu)
Covariance has units equal to the product of the variables’ units. Changing a measurement from meters to centimeters multiplies its covariance with an unchanged second variable by 100. Consequently, an unstandardized covariance cannot by itself provide a scale-independent assessment of association. (ocw.mit.edu)
Algebraic properties
Covariance is symmetric and reduces to variance when both arguments are the same:
For constants ,
Thus, adding constants does not change covariance, while multiplication rescales it and can reverse its sign. Covariance is also additive in each argument:
These properties explain its usefulness for studying linear combinations of random quantities. In particular,
The covariance term accounts for the contribution of joint variation to the variance of a sum. (ocw.mit.edu)
Covariance, correlation, and independence
When both standard deviations are finite and positive, Pearson’s population correlation is
This normalization removes measurement units. The Cauchy–Schwarz inequality gives
so correlation lies between and . Equality corresponds to an exact affine relationship almost surely, provided both variances are positive. If either variable has zero variance, covariance is zero but this correlation formula is undefined. (online.stat.psu.edu)
Statistical independence implies zero covariance when the required expectations exist. The converse generally fails. As an illustrative calculation, let take the values , each with probability , and let . Then
giving zero covariance even though is determined by . This is a nonlinear dependence that covariance does not detect. The example illustrates the general distinction between independence and uncorrelatedness. (ocw.mit.edu)
An important exception occurs for jointly Gaussian variables: within a multivariate normal distribution, zero covariance between two components implies their independence. Merely having individually normal marginal distributions is insufficient for this conclusion. (probabilitycourse.com)
Estimation from observations
For paired observations , the usual sample covariance is
where and are sample means. Under independent, identically distributed sampling of the pairs, with finite second moments, this estimator has zero bias for population covariance. The denominator applies the same Bessel’s correction used in sample variance. (online.stat.psu.edu)
Dividing the centered sum by instead gives the covariance of the empirical distribution assigning equal probability to every observed pair. This convention is useful descriptively, but under the sampling assumptions above its expectation is times the population covariance. The denominator must therefore be specified when reporting or comparing calculations. (stats.libretexts.org)
Covariance matrices and applications
For a random vector , the covariance matrix collects all pairwise covariances:
Its diagonal entries are variances, and its off-diagonal entries are covariances. This matrix is symmetric and positive semidefinite because, for every real vector ,
For a fixed matrix , a transformed vector has covariance . (web.mit.edu)
In simple linear regression with an intercept, the ordinary least squares slope is , provided the predictor’s sample variance is positive. Covariance therefore determines the slope’s sign and contributes directly to its magnitude. (online.stat.psu.edu)
In principal component analysis, the covariance matrix’s eigenvectors identify mutually orthogonal directions of variation, while the corresponding eigenvalues give projected variances. This supports dimensionality reduction in machine learning. Because covariance depends on scale, PCA based on raw covariance can differ substantially from PCA based on standardized variables or a correlation matrix. (online.stat.psu.edu)