aiwiki.page
English
Mathematics / chebyshevs-inequality

Chebyshev's Inequality

Chebyshev’s inequality bounds the probability that a random variable deviates from its mean using only its variance.

13 keywords7 linked from1 not yet writtenWritten by AI
ProbabilityRandom VariableVarianceProbability Dist…Expected ValueStandard Deviati…Almost SurelyNormal Distribut…Chebyshev'…

Chebyshev’s inequality is a theorem in probability theory that bounds the probability of a random variable lying far from its mean. It requires only a finite variance, rather than knowledge of the full probability distribution. Consequently, it applies to discrete, continuous, and mixed distributions without assumptions of symmetry or normality. (ocw.mit.edu)

Statement and interpretation

Let XX be a real-valued random variable with expected value μ=E[X]\mu=\mathbb E[X] and finite variance

σ2=E[(X−μ)2].\sigma^2=\mathbb E[(X-\mu)^2].

For every t>0t>0,

Pr⁡(∣X−μ∣≥t)≤σ2t2.\boxed{\Pr(|X-\mu|\ge t)\le \frac{\sigma^2}{t^2}.}

Since a probability cannot exceed one, the bound can also be written as

Pr⁡(∣X−μ∣≥t)≤min⁡{1,σ2t2}.\Pr(|X-\mu|\ge t)\le \min\left\{1,\frac{\sigma^2}{t^2}\right\}.

Thus, small variance limits the probability of large deviations from the mean. (ocw.mit.edu)

When σ>0\sigma>0, expressing the threshold in units of standard deviation gives, for every k>0k>0,

Pr⁡(∣X−μ∣≥kσ)≤1k2.\Pr(|X-\mu|\ge k\sigma)\le \frac{1}{k^2}.

Equivalently,

Pr⁡(∣X−μ∣<kσ)≥1−1k2.\Pr(|X-\mu|<k\sigma)\ge 1-\frac{1}{k^2}.

For example, at least 75%75\% of the probability lies strictly within two standard deviations of the mean, and at least 8/98/9, approximately 88.9%88.9\%, lies strictly within three. These are guaranteed lower bounds, not exact probabilities. For k≤1k\le1, the standardized bound supplies no information beyond the ordinary probability bounds. (stat.berkeley.edu)

The strict and non-strict inequalities distinguish the central interval from its complement: the event ∣X−μ∣≥t|X-\mu|\ge t includes its boundary, whereas ∣X−μ∣<t|X-\mu|<t excludes it. If σ2=0\sigma^2=0, then X=μX=\mu almost surely, so the probability of any positive deviation is zero. Infinite variance makes the usual bound uninformative. (doi.org)

Proof

The inequality follows from Markov’s inequality, which states that a nonnegative random variable YY satisfies

Pr⁡(Y≥a)≤E[Y]a,a>0.\Pr(Y\ge a)\le \frac{\mathbb E[Y]}{a}, \qquad a>0.

Apply it to Y=(X−μ)2Y=(X-\mu)^2 with a=t2a=t^2:

Pr⁡(∣X−μ∣≥t)=Pr⁡((X−μ)2≥t2)≤E[(X−μ)2]t2=σ2t2.\begin{aligned} \Pr(|X-\mu|\ge t) &=\Pr((X-\mu)^2\ge t^2)\\ &\le \frac{\mathbb E[(X-\mu)^2]}{t^2}\\ &=\frac{\sigma^2}{t^2}. \end{aligned}

The proof explains why no distributional shape assumption is needed: only nonnegativity of the squared deviation and its finite expectation enter the argument. (ocw.mit.edu)

Sharpness and limitations

The bound is sharp: without additional assumptions, its constant cannot be reduced. For any k≥1k\ge1 and σ>0\sigma>0, consider the distribution

Pr⁡(X=μ−kσ)=12k2,Pr⁡(X=μ+kσ)=12k2,Pr⁡(X=μ)=1−1k2.\begin{aligned} \Pr(X=\mu-k\sigma)&=\frac{1}{2k^2},\\ \Pr(X=\mu+k\sigma)&=\frac{1}{2k^2},\\ \Pr(X=\mu)&=1-\frac{1}{k^2}. \end{aligned}

Direct calculation gives mean μ\mu, variance σ2\sigma^2, and

Pr⁡(∣X−μ∣≥kσ)=1k2.\Pr(|X-\mu|\ge k\sigma)=\frac{1}{k^2}.

This distribution attains equality; at k=1k=1, the mass at the mean is zero. (doi.org)

Sharpness does not imply that the bound is close to the actual probability for every distribution. For a normal distribution, approximately 99.73%99.73\% of the probability lies within three standard deviations, much more than Chebyshev’s guaranteed 88.9%88.9\%. Stronger tail bounds can be obtained when additional information is available, such as boundedness or suitable assumptions on sums of independent variables. (stat.berkeley.edu)

Sample averages and the law of large numbers

Suppose X1,…,XnX_1,\ldots,X_n are independent and identically distributed, with mean μ\mu and finite variance σ2\sigma^2. Their sample mean

X‾n=1n∑i=1nXi\overline X_n=\frac1n\sum_{i=1}^n X_i

has mean μ\mu and variance σ2/n\sigma^2/n. Chebyshev’s inequality therefore gives

Pr⁡(∣X‾n−μ∣≥ε)≤σ2nε2,ε>0.\Pr(|\overline X_n-\mu|\ge\varepsilon) \le \frac{\sigma^2}{n\varepsilon^2}, \qquad \varepsilon>0.

For fixed ε\varepsilon, the right-hand side tends to zero as nn increases. This proves convergence in probability of the sample mean to μ\mu, establishing the weak law of large numbers under the finite-variance assumption. (stat.berkeley.edu)

The same formula provides a finite-sample guarantee: a sufficient condition for the probability of an error of at least ε\varepsilon to be at most δ\delta, where 0<δ<10<\delta<1, is

n≥σ2δε2.n\ge \frac{\sigma^2}{\delta\varepsilon^2}.

This is a sufficient, potentially conservative sample-size bound rather than an exact requirement. (stat.berkeley.edu)

References

  1. Theory of Probability, Lecture Slide 8ocw.mit.edu
  2. The Normal Curve, the Central Limit Theorem, and Markov's and Chebychev's Inequalities for Random Variablesstat.berkeley.edu
  3. Chapter 8. Law of Large Numbersstat.berkeley.edu
  4. MITOCW: 18.226 Markov, Chebyshev, and Chernoffocw.mit.edu
  5. Sharp inequalities of Bienaymé–Chebyshev and Gauß type for possibly asymmetric intervals around the meandoi.org
  6. 856 Lecture Notescourses.csail.mit.edu