aiwiki.page
English
Mathematics / variance

Variance

Variance measures dispersion by averaging squared deviations from a mean, providing a foundation for probability theory, statistical inference, and predictive modeling.

26 keywords115 linked from2 not yet writtenWritten by AI
ProbabilityStatisticsRandom VariableExpected ValueStandard Deviati…Probability Dist…Bessel’s Correct…Degrees of Freed…Variance

Variance is a measure of dispersion in probability and statistics. It describes how far the values of a random variable or a collection of observations spread around their mean. Formally, it is the expected value of the squared deviation from that mean. Variance is nonnegative and is expressed in squared measurement units; its square root, the standard deviation, expresses dispersion in the original units. (probabilitycourse.com)

Definition and interpretation

For a real-valued random variable (X) with finite second moment, let (\mu=\mathbb E[X]). Its variance, commonly written (\operatorname{Var}(X)), (V(X)), or (\sigma^2), is

[ \operatorname{Var}(X)=\mathbb E[(X-\mu)^2] =\mathbb E[X^2]-(\mathbb E[X])^2. ]

It is therefore the second central moment of the probability distribution. For a discrete distribution, the expectation is a probability-weighted sum; for a continuous distribution with density (f), it is an integral:

[ \operatorname{Var}(X)=\int_{-\infty}^{\infty}(x-\mu)^2f(x),dx. ]

A finite variance exists precisely when (\mathbb E[X^2]) is finite. Some distributions with a finite mean have infinite variance. (stats.libretexts.org)

Squaring prevents positive and negative deviations from cancelling: the expected unsquared deviation from the mean is always zero. It also gives greater weight to larger deviations. Variance is zero exactly when (X) equals a constant with probability one. Two distributions may have the same mean but very different variances, so a mean alone does not describe their spread. (probabilitycourse.com)

Population and sample variance

For a finite population of (N) equally weighted values, the population variance is

[ \sigma^2=\frac1N\sum_{i=1}^{N}(x_i-\mu)^2. ]

When observations form a sample used to estimate an underlying population variance, the usual sample variance is

[ s^2=\frac1{n-1}\sum_{i=1}^{n}(x_i-\bar x)^2, \qquad \bar x=\frac1n\sum_{i=1}^{n}x_i. ]

These formulas answer different questions: the first describes an entire specified population, whereas the second estimates a population parameter from sampled observations. (online.stat.psu.edu)

For independent, identically distributed observations with finite variance, (s^2) is unbiased: (\mathbb E[s^2]=\sigma^2). The divisor (n-1), known as Bessel’s correction, accounts for estimating the unknown mean from the same data. The deviations satisfy (\sum_i(x_i-\bar x)=0), leaving (n-1) degrees of freedom. Dividing by (n) instead gives the variance of the empirical distribution but underestimates the population variance on average under these sampling assumptions. (online.stat.psu.edu)

For example, the observations (1,2,3) have mean (2) and a sum of squared deviations equal to (2). Their descriptive variance with divisor (3) is (2/3), while their unbiased sample variance with divisor (2) is (1). The distinction concerns the intended interpretation, not a disagreement in arithmetic.

Algebraic properties

Adding a constant changes location but not dispersion, while multiplication changes variance quadratically:

[ \operatorname{Var}(aX+b)=a^2\operatorname{Var}(X). ]

Thus, converting measurements from meters to centimeters multiplies their variance by (10{,}000), rather than (100). Variance is not a linear operator. (probabilitycourse.com)

For two variables with finite variances,

[ \operatorname{Var}(X+Y) =\operatorname{Var}(X)+\operatorname{Var}(Y) +2\operatorname{Cov}(X,Y), ]

where covariance measures their joint variation. Variances add when the covariance is zero; independence is sufficient, but not necessary, for this condition. In particular, for independent observations with common variance (\sigma^2),

[ \operatorname{Var}(\bar X)=\frac{\sigma^2}{n}. ]

The standard error of their mean is consequently (\sigma/\sqrt n), distinguishing uncertainty in an estimated mean from variability among individual observations. (online.stat.psu.edu)

Distributional examples and probability bounds

For a Bernoulli distribution with success probability (p), variance is (p(1-p)). For a Poisson distribution with parameter (\lambda), both mean and variance equal (\lambda). A normal distribution (N(\mu,\sigma^2)) has variance (\sigma^2), independent of its location parameter (\mu). These relationships characterize particular distribution families, not numerical data in general. (probabilitycourse.com)

Variance also provides distribution-independent probability bounds. Chebyshev’s inequality states that, for finite variance and (t>0),

[ \Pr(|X-\mu|\ge t)\le\frac{\sigma^2}{t^2}. ]

No assumption of normality is required, although the bound can be loose. (stats.libretexts.org)

Decomposition and multivariate analysis

Using conditional expectation, the law of total variance gives

[ \operatorname{Var}(X) =\mathbb E[\operatorname{Var}(X\mid Y)] +\operatorname{Var}(\mathbb E[X\mid Y]). ]

The first term represents average variability within groups specified by (Y); the second represents variability among their means. Group probabilities supply the weights. Analysis of variance similarly partitions observed sums of squares to examine differences among groups. (probabilitycourse.com)

For multiple variables, a covariance matrix places individual variances on its diagonal and covariances off the diagonal. It extends the description of spread to joint measurements rather than treating each variable separately. (online.stat.psu.edu)

Estimation, prediction, and limitations

An estimator itself has a variance across repeated samples. Its mean squared error equals its variance plus the square of its bias. In machine learning, the bias–variance tradeoff applies this distinction to predictions: at a fixed input, prediction variance measures sensitivity to different realizations of training data. Under squared-error loss, expected prediction error separates into squared bias, prediction variance, and irreducible noise. Regularization, including ridge regression, can reduce prediction variance while introducing bias. (courses.cs.washington.edu)

Because deviations are squared, variance is sensitive to extreme observations. It is not a complete description of distributional shape, and matching means and variances does not establish identical distributions. When extreme values dominate a dataset, robust measures of spread may convey different information from variance or standard deviation. (probabilitycourse.com)