In statistics, a consistent estimator is an estimator whose error converges to zero in probability as the sample size increases without bound. For any fixed positive tolerance, the probability that the estimate differs from its target by more than that tolerance tends to zero. Consistency describes large-sample behavior, rather than guaranteeing accuracy for any particular finite sample. Strictly, it is a property of a sequence of estimators indexed by sample size. (ocw.mit.edu)
Formal definition
Let be observations from a statistical model with parameter , and let
estimate . The sequence is consistent at if, for every ,
Here denotes probability under the model with parameter . This is written
using the notation for convergence in probability. An estimator is consistent on if this condition holds for every fixed . When the target is a function , the definition replaces by . (ocw.mit.edu)
For vector-valued parameters, absolute error is replaced by a norm, or more generally by distance in a metric space. Consistency concerns concentration near the target; it does not require the realized error to decrease whenever another observation is added. (ocw.mit.edu)
Forms of consistency
Weak consistency is the convergence-in-probability definition above. Unless otherwise specified, “consistent” ordinarily means weakly consistent. Strong consistency instead requires convergence almost surely:
Strong consistency implies weak consistency, but the converse does not hold in general. (statlect.com)
Mean-square consistency requires
It implies weak consistency because
This is stronger than merely requiring small errors with high probability: rare, very large errors can prevent mean squared error from approaching zero. (bpb-us-w2.wpmucdn.com)
Uniform consistency strengthens the quantifier over parameter values:
Ordinary consistency applies separately to each fixed parameter; uniform consistency controls the worst-case error probability over the entire parameter space. This is a different distinction from weak versus strong convergence. (ocw.mit.edu)
Consistency, bias, and mean squared error
Bias concerns the difference between an estimator’s expected value and its target:
Unbiasedness requires this difference to be zero at each sample size. Consistency instead concerns a limit of error probabilities. Neither property, by itself, implies the other. (bpb-us-w2.wpmucdn.com)
For example, suppose are independent observations with mean and positive finite variance. The estimator is unbiased for , but its distribution does not change with ; it therefore does not become concentrated at . This illustrates why unbiasedness alone cannot establish consistency. (bpb-us-w2.wpmucdn.com)
When the second moment exists, mean squared error satisfies
Consequently, bias tending to zero together with variance tending to zero is a sufficient condition for consistency. It is not a necessary condition for weak consistency: convergence in probability alone need not imply convergence of expectations or second moments. (bpb-us-w2.wpmucdn.com)
Basic examples
Sample mean
Let be independent and identically distributed, with and mean . The sample mean
is consistent for by the law of large numbers. Finite variance is not necessary for this conclusion. (ocw.mit.edu)
If , there is a direct proof using Chebyshev’s inequality:
In this case the sample mean is also mean-square consistent, since its mean squared error is . (bpb-us-w2.wpmucdn.com)
A biased but consistent endpoint estimator
For independent observations from the uniform distribution on , where , consider the sample maximum . Its expectation is
so it is biased downward. Nevertheless, for ,
Thus the maximum is consistent despite its finite-sample bias. (ocw.mit.edu)
Establishing consistency
Many proofs combine a law of large numbers with a continuous transformation. If is consistent for , and is a function continuous at , the continuous mapping theorem gives
This underlies many plug-in estimators: estimated moments or probabilities are substituted into a continuous formula for the target. (ocw.mit.edu)
Estimators defined by optimization require additional reasoning. An M-estimator minimizes or maximizes a sample criterion. A common sufficient framework is that the criterion converges uniformly to a deterministic population criterion, whose optimum is uniquely separated from values outside every neighborhood of the target. Suitable approximate optimizers can also be consistent. Pointwise convergence of the criterion alone generally does not justify convergence of its optimizer. (ocw.mit.edu)
Maximum likelihood estimation fits this framework through the average log-likelihood. Its consistency requires appropriate model and regularity conditions; maximization of a likelihood is not, by itself, a guarantee. Identification is fundamental: distinct target values must be distinguishable through the observational distribution. (ocw.mit.edu)
In linear regression, the ordinary least squares estimator is consistent under suitable laws of large numbers, a nonsingular limiting regressor second-moment matrix, and population orthogonality between regressors and errors. These assumptions are separate from assumptions used to establish its asymptotic normality. (statlect.com)
Rates and limits of interpretation
Consistency does not specify a convergence rate or a limiting distribution for scaled errors. For an independent, identically distributed sample with positive finite variance, the central limit theorem supplies the stronger result
Such distributional results, together with consistent variance estimation, support asymptotic confidence intervals and hypothesis tests. Consistency alone does not provide their calibration. (ocw.mit.edu)
A consistency claim is conditional on its sampling model and assumptions. Dependence, lack of identification, or failure of required moment conditions can invalidate a particular proof. For regression, increasing the sample size does not replace the orthogonality condition needed for consistency of ordinary least squares. Finite-sample performance remains a separate question: an estimator can be consistent while having substantial bias, variability, or computational cost at practically available sample sizes. (statlect.com)
References
- M-estimators and their consistencyocw.mit.edu
- More about Estimators and Mathematical Digressionsstat135.berkeley.edu
- Mathematical Statistics, Lecture 16 Asymptotics: Consistency and Delta Methodocw.mit.edu
- Point estimation: Theory and examplesstatlect.com
- Lecture: Consistencybpb-us-w2.wpmucdn.com
- MITOCW: 14.310x Lecture 12 transcriptocw.mit.edu
- Consistency of the Maximum Likelihood Estimatorstat135.berkeley.edu
- Lecture 7: Maximum Likelihood Estimationocw.mit.edu
- Properties of the OLS estimator: Consistency, asymptotic normalitystatlect.com