Standard deviation is a measure of dispersion in statistics and probability, describing how widely observations or possible values are spread around their mean. It is the nonnegative square root of variance. Because it has the same units as the measured quantity, it expresses variability on the original measurement scale rather than in squared units. Population standard deviation is conventionally denoted by , while sample standard deviation is commonly denoted by . (itl.nist.gov)
Definition
For a random variable with expected value and finite second moment, standard deviation is defined by
The expression inside the square root is the second central moment. Thus, standard deviation describes a probability distribution, not merely a collection of observations. For discrete distributions, the expectation is calculated by weighting squared deviations by their probabilities; for continuous distributions, it is calculated using an integral against the probability density. (online.stat.psu.edu)
For a complete finite population of equally weighted values, the corresponding formula is
It is a root-mean-square deviation, not an average absolute distance. Squaring makes positive and negative deviations contribute positively and gives larger deviations disproportionately greater weight. (online.stat.psu.edu)
Sample estimation
When observations are used to estimate variability in a larger population, the conventional sample standard deviation is
for . The denominator , called Bessel’s correction, reflects the loss of one degree of freedom when the population mean is estimated from the same observations: the deviations from the sample mean must sum to zero. (online.stat.psu.edu)
For independent, identically distributed observations with finite variance, is an unbiased estimator of population variance. However, taking its square root does not preserve unbiasedness: generally has a downward estimator bias for . Unbiased variance estimation and unbiased standard-deviation estimation are therefore distinct problems. (online.stat.psu.edu)
Dividing by instead describes the dispersion of the observed values treated as an equally weighted empirical population. The choice between and concerns the intended statistical interpretation, not simply whether the dataset is large or small. (online.stat.psu.edu)
Worked example
Consider the values . Their mean is , and their squared deviations sum to
Applying the population formula gives
Applying the conventional sample formula to the same values gives
These are calculations under the two definitions above. The difference arises entirely from the denominator. Neither result means that every observation is that distance from the mean, or that the standard deviation equals the dataset’s range.
Mathematical properties
Standard deviation is nonnegative. For a finite dataset, it equals zero precisely when every value is identical; for a random variable with finite variance, it equals zero when the variable is constant with probability one. Adding a constant changes location but not dispersion. Multiplication by a constant changes standard deviation by that constant’s absolute magnitude:
Consequently, a change of measurement units rescales standard deviation in the same way as the observations. (online.stat.psu.edu)
For two variables with finite variances,
The covariance term represents their joint variation. Under statistical independence, it is zero, so variances add; standard deviations ordinarily do not. (online.stat.psu.edu)
Distributional interpretation
For a normal distribution, approximately 68.27% of probability lies within one standard deviation of the mean, 95.45% within two, and 99.73% within three. These percentages depend on normality and are not universal properties of standard deviation. (itl.nist.gov)
A broader statement follows from Chebyshev’s inequality. For any distribution with finite, positive standard deviation and ,
Thus, at least 75% of probability lies within two standard deviations, without requiring symmetry or normality. This bound is generally much less specific than the normal-distribution percentages. (nvlpubs.nist.gov)
A z-score, , expresses an observation’s signed distance from the mean in standard-deviation units. Standardization changes location and scale; it does not by itself make a nonnormal distribution normal. (itl.nist.gov)
Standard error, limitations, and computation
Standard deviation describes variability of observations, whereas standard error describes variability of an estimator across repeated samples. For independent, identically distributed observations, the sample mean has standard error , commonly estimated by . This distinction is important when interpreting a confidence interval: uncertainty about a mean is not the same as dispersion among individual observations. (online.stat.psu.edu)
Because squared deviations emphasize extreme observations, standard deviation is sensitive to outliers and heavy tails. The interquartile range and median absolute deviation provide alternative descriptions of spread with less sensitivity to extreme values. Some distributions lack a finite population standard deviation, even though any finite sample of finite observations has a calculable sample standard deviation. (itl.nist.gov)
Numerical implementation also matters. Computing variance by subtracting two large, nearly equal quantities can lose precision. Calculating deviations from the mean before squaring and summing avoids this particular cancellation problem, which is especially important when observations have a large common offset but relatively little variation. (nist.gov)