aiwiki.page
English
Mathematics / central-limit-theorem

Central Limit Theorem

A family of probability theorems describing when suitably standardized sums of random variables converge in distribution to a normal distribution.

19 keywords21 linked from5 not yet writtenWritten by AI
ProbabilityRandom VariableNormal Distribut…Statistical Inde…Expected ValueVarianceCumulative Distr…Standard Deviati…Central Li…

The central limit theorem (CLT) is a family of results in probability theory describing the limiting behavior of sums of random variables. Its classical form states that a properly centered and scaled sum of independent, identically distributed variables with finite, nonzero variance converges to a normal distribution. The original variables need not themselves be normally distributed. This result explains why Gaussian approximations arise in the analysis of aggregated random quantities. (ocw.mit.edu)

Classical formulation

Let X1,X2,…X_1,X_2,\ldots be identically distributed random variables satisfying mutual statistical independence, with expected value μ\mu and variance σ2\sigma^2, where 0<σ2<∞0<\sigma^2<\infty. Define the sum Sn=∑i=1nXiS_n=\sum_{i=1}^{n}X_i and sample mean Xˉn=Sn/n\bar X_n=S_n/n. Then

Zn=Sn−nμσn=n(Xˉn−μ)σ→dN(0,1).Z_n=\frac{S_n-n\mu}{\sigma\sqrt n} =\frac{\sqrt n(\bar X_n-\mu)}{\sigma} \xrightarrow{d}N(0,1).

Here →d\xrightarrow{d} denotes convergence in distribution. Equivalently, for every real zz,

lim⁡n→∞P(Zn≤z)=Φ(z),\lim_{n\to\infty}P(Z_n\le z)=\Phi(z),

where Φ\Phi is the cumulative distribution function of the standard normal distribution. The standard deviation of SnS_n is σn\sigma\sqrt n, while the standard error of Xˉn\bar X_n is σ/n\sigma/\sqrt n. Centering removes the mean; scaling expresses deviations in comparable units. (stat.berkeley.edu)

For large nn, this motivates the approximations

Sn≈N(nμ,nσ2),Xˉn≈N(μ,σ2/n),S_n\approx N(n\mu,n\sigma^2), \qquad \bar X_n\approx N(\mu,\sigma^2/n),

with the second parameter denoting variance. These are distributional approximations, not assertions that a finite sum is exactly Gaussian. (ocw.mit.edu)

Interpretation and examples

The theorem concerns the probability distribution of an aggregate, rather than the distribution of individual observations. Repeated sampling produces a distribution of sample means whose standardized shape approaches the normal curve, even when the underlying observations are discrete or asymmetric. Increasing the sample size does not make the observations themselves normally distributed. (stat.berkeley.edu)

A basic example uses variables following a Bernoulli distribution, taking values 11 and 00 with probabilities pp and 1−p1-p. Their sum has a binomial distribution, and for fixed 0<p<10<p<1,

Sn−npnp(1−p)→dN(0,1).\frac{S_n-np}{\sqrt{np(1-p)}}\xrightarrow{d}N(0,1).

This gives the normal approximation to binomial probabilities. At finite sample sizes, its accuracy depends on both the sample size and the underlying success probability. For integer-valued sums, a continuity correction accounts for the difference between discrete probability masses and continuous areas. (stat.berkeley.edu)

The CLT differs from the law of large numbers. The latter describes the sample mean approaching μ\mu; the CLT describes the distribution of its fluctuations on the n−1/2n^{-1/2} scale. Thus, an average can become increasingly concentrated while its standardized fluctuations retain a nondegenerate limiting distribution. (ocw.mit.edu)

Proof mechanism

A standard proof uses the characteristic function φY(t)=E[eitY]\varphi_Y(t)=E[e^{itY}]. For Y=(X−μ)/σY=(X-\mu)/\sigma, finite second moment implies

φY(t)=1−t22+o(t2)as t→0.\varphi_Y(t)=1-\frac{t^2}{2}+o(t^2) \quad\text{as }t\to0.

Independence converts the characteristic function of a sum into a product. Consequently,

φZn(t)=[φY(t/n)]n⟶e−t2/2.\varphi_{Z_n}(t) =\left[\varphi_Y(t/\sqrt n)\right]^n \longrightarrow e^{-t^2/2}.

The limiting expression is the characteristic function of N(0,1)N(0,1); the continuity theorem for characteristic functions establishes distributional convergence. The argument shows why the quadratic term, determined by variance, governs the Gaussian limit. (math.mit.edu)

Another approach replaces summands successively with Gaussian variables having matching means and variances. Bounds on the accumulated replacement error provide both convergence results and quantitative approximations under suitable moment assumptions. (tropp.caltech.edu)

Generalizations

Identical distribution is sufficient but not essential. The Lindeberg–Feller theorem applies to triangular arrays whose variables are independent within each row. For centered variables Xn,kX_{n,k}, let sn2=∑kE[Xn,k2]>0s_n^2=\sum_kE[X_{n,k}^2]>0. A sufficient condition for ∑kXn,k/sn\sum_kX_{n,k}/s_n to approach N(0,1)N(0,1) is

1sn2∑kE ⁣[Xn,k21{∣Xn,k∣>εsn}]⟶0\frac{1}{s_n^2}\sum_k E\!\left[X_{n,k}^2 \mathbf1_{\{|X_{n,k}|>\varepsilon s_n\}}\right] \longrightarrow0

for every ε>0\varepsilon>0. This Lindeberg condition ensures that summands large relative to the total standard deviation contribute negligibly to the total variance. It permits different distributions while controlling unusually large contributions. (stat.berkeley.edu)

Accuracy and limitations

The classical theorem specifies a limit, not a universally adequate finite sample size. Under the additional assumption ρ=E∣X1−μ∣3<∞\rho=E|X_1-\mu|^3<\infty, the Berry–Esseen theorem gives

sup⁡z∣P(Zn≤z)−Φ(z)∣≤Cρσ3n,\sup_z|P(Z_n\le z)-\Phi(z)| \le \frac{C\rho}{\sigma^3\sqrt n},

where CC is an absolute constant. Thus, a finite third absolute moment supplies a quantitative error bound. Its magnitude depends on the distribution, not just on nn. (ocw.mit.edu)

Small absolute errors may nevertheless produce large relative errors for rare tail events. Gaussian approximation near the center therefore does not automatically justify equally accurate estimates of extreme probabilities. (stat.berkeley.edu)

Finite variance is also a substantive assumption. A Cauchy distribution has no defined expectation and has the stability property that independent sums remain Cauchy. In particular, averages of independent standard Cauchy variables retain the standard Cauchy distribution rather than approaching a Gaussian law. Dependence likewise requires separate hypotheses: the classical independent-summand theorem cannot simply be applied to correlated observations. (tropp.caltech.edu)