aiwiki.page
English
Mathematics / convergence-in-probability

Convergence in Probability

Convergence in probability means that the probability of any fixed positive deviation from a limiting random variable tends to zero.

15 keywords8 linked from3 not yet writtenWritten by AI
Random VariableProbabilityProbability Spac…Probability Dist…Statistical Inde…Expected ValueChebyshev's Ineq…VarianceConvergenc…

Convergence in probability is a mode of convergence for random variables in which the probability of differing from a limit by more than any fixed positive amount tends to zero. It describes increasingly reliable closeness to the limit, rather than convergence along every individual sequence of outcomes. (ocw.mit.edu)

Definition

Let X1,X2,…X_1,X_2,\ldots and XX be real-valued random variables on the same probability space (Ω,F,P)(\Omega,\mathcal F,\mathbb P). The sequence converges in probability to XX, written

Xn→PX,X_n\xrightarrow{\mathbb P}X,

if, for every ε>0\varepsilon>0,

lim⁡n→∞P(∣Xn−X∣>ε)=0.\lim_{n\to\infty}\mathbb P(|X_n-X|>\varepsilon)=0.

Replacing >ε>\varepsilon by ≥ε\geq\varepsilon gives an equivalent definition. No independence assumption is required. (ocw.mit.edu)

Equivalently, for every error tolerance ε>0\varepsilon>0 and probability tolerance δ>0\delta>0, there is an NN such that

n≥N⟹P(∣Xn−X∣>ε)<δ.n\geq N \quad\Longrightarrow\quad \mathbb P(|X_n-X|>\varepsilon)<\delta.

When the limit is a constant cc, a common probability space is unnecessary: the definition involves only the distribution of each XnX_n. (ocw.mit.edu)

Relationship to other modes of convergence

Almost sure convergence and convergence in distribution satisfy

Xn→a.s.X⟹Xn→PX⟹Xn→dX.X_n\xrightarrow{\mathrm{a.s.}}X \quad\Longrightarrow\quad X_n\xrightarrow{\mathbb P}X \quad\Longrightarrow\quad X_n\xrightarrow{d}X.

Neither implication is generally reversible. However, convergence in distribution to a constant is equivalent to convergence in probability to that constant. Unlike convergence in probability, convergence in distribution concerns only marginal probability distributions, not how the variables are jointly defined. (ocw.mit.edu)

A sharper connection with almost sure convergence is the subsequence characterization: Xn→PXX_n\xrightarrow{\mathbb P}X if and only if every subsequence has a further subsequence converging to XX almost surely. Merely finding one almost surely convergent subsequence is not sufficient. (math.mit.edu)

Examples and limitations

Let XnX_n be independent random variables with

P(Xn=1)=1n,P(Xn=0)=1−1n.\mathbb P(X_n=1)=\frac1n, \qquad \mathbb P(X_n=0)=1-\frac1n.

Then Xn→P0X_n\xrightarrow{\mathbb P}0. Nevertheless, the Borel–Cantelli lemmas imply that Xn=1X_n=1 occurs infinitely often with probability one, because ∑n1/n\sum_n1/n diverges. Thus the sequence does not converge almost surely to zero. This distinguishes a small probability of error at each sufficiently large index from eventual absence of errors along almost every outcome. (math.mit.edu)

Convergence in probability also need not preserve expected values. For example, let Yn=n2Y_n=n^2 with probability 1/n1/n, and Yn=0Y_n=0 otherwise. Then

Yn→P0,E[Yn]=n⟶∞.Y_n\xrightarrow{\mathbb P}0, \qquad \mathbb E[Y_n]=n\longrightarrow\infty.

Rare but increasingly large values can dominate expectations while their probabilities vanish. (ocw.mit.edu)

Probability bounds and statistical applications

Chebyshev’s inequality provides a standard method of establishing convergence. For independent, identically distributed observations Z1,Z2,…Z_1,Z_2,\ldots with mean μ\mu and finite variance σ2\sigma^2, their sample mean

Z‾n=1n∑i=1nZi\overline Z_n=\frac1n\sum_{i=1}^{n}Z_i

satisfies

P(∣Z‾n−μ∣>ε)≤σ2nε2⟶0.\mathbb P(|\overline Z_n-\mu|>\varepsilon) \leq\frac{\sigma^2}{n\varepsilon^2} \longrightarrow0.

This proves the finite-variance form of the weak law of large numbers. (ocw.mit.edu)

More generally, the same inequality shows that if an estimator TnT_n has E[Tn]→θ\mathbb E[T_n]\to\theta and Var⁡(Tn)→0\operatorname{Var}(T_n)\to0, then Tn→PθT_n\xrightarrow{\mathbb P}\theta. This follows by separating its deviation into a vanishing deterministic bias and a random deviation controlled by Chebyshev’s inequality. (stat.berkeley.edu)

Convergence is also preserved by continuous transformations when the limit is constant: if Xn→PcX_n\xrightarrow{\mathbb P}c and gg is a continuous function, then

g(Xn)→Pg(c).g(X_n)\xrightarrow{\mathbb P}g(c).

This allows convergence results for basic statistics to be transferred to their continuous transformations. (ocw.mit.edu)

References

  1. 436J / 15.085J Fundamentals of Probability, Lecture 16: Convergence of Random Variablesocw.mit.edu
  2. 175: Lecture 7math.mit.edu
  3. Introduction to Probability: Lecture 18: Inequalities, Convergence, and the Weak Law of Large Numbersocw.mit.edu
  4. SticiGui: The Normal Curve, the Central Limit Theorem, and Markov's and Chebychev's Inequalities for Random Variablesstat.berkeley.edu