Convergence in probability is a mode of convergence for random variables in which the probability of differing from a limit by more than any fixed positive amount tends to zero. It describes increasingly reliable closeness to the limit, rather than convergence along every individual sequence of outcomes. (ocw.mit.edu)
Definition
Let and be real-valued random variables on the same probability space . The sequence converges in probability to , written
if, for every ,
Replacing by gives an equivalent definition. No independence assumption is required. (ocw.mit.edu)
Equivalently, for every error tolerance and probability tolerance , there is an such that
When the limit is a constant , a common probability space is unnecessary: the definition involves only the distribution of each . (ocw.mit.edu)
Relationship to other modes of convergence
Almost sure convergence and convergence in distribution satisfy
Neither implication is generally reversible. However, convergence in distribution to a constant is equivalent to convergence in probability to that constant. Unlike convergence in probability, convergence in distribution concerns only marginal probability distributions, not how the variables are jointly defined. (ocw.mit.edu)
A sharper connection with almost sure convergence is the subsequence characterization: if and only if every subsequence has a further subsequence converging to almost surely. Merely finding one almost surely convergent subsequence is not sufficient. (math.mit.edu)
Examples and limitations
Let be independent random variables with
Then . Nevertheless, the Borel–Cantelli lemmas imply that occurs infinitely often with probability one, because diverges. Thus the sequence does not converge almost surely to zero. This distinguishes a small probability of error at each sufficiently large index from eventual absence of errors along almost every outcome. (math.mit.edu)
Convergence in probability also need not preserve expected values. For example, let with probability , and otherwise. Then
Rare but increasingly large values can dominate expectations while their probabilities vanish. (ocw.mit.edu)
Probability bounds and statistical applications
Chebyshev’s inequality provides a standard method of establishing convergence. For independent, identically distributed observations with mean and finite variance , their sample mean
satisfies
This proves the finite-variance form of the weak law of large numbers. (ocw.mit.edu)
More generally, the same inequality shows that if an estimator has and , then . This follows by separating its deviation into a vanishing deterministic bias and a random deviation controlled by Chebyshev’s inequality. (stat.berkeley.edu)
Convergence is also preserved by continuous transformations when the limit is constant: if and is a continuous function, then
This allows convergence results for basic statistics to be transferred to their continuous transformations. (ocw.mit.edu)
References
- 436J / 15.085J Fundamentals of Probability, Lecture 16: Convergence of Random Variablesocw.mit.edu
- 175: Lecture 7math.mit.edu
- Introduction to Probability: Lecture 18: Inequalities, Convergence, and the Weak Law of Large Numbersocw.mit.edu
- SticiGui: The Normal Curve, the Central Limit Theorem, and Markov's and Chebychev's Inequalities for Random Variablesstat.berkeley.edu