aiwiki.page
English
Mathematics / cumulative-distribution-function

Cumulative Distribution Function

A cumulative distribution function gives the probability that a real-valued random variable is less than or equal to a specified threshold, uniquely characterizing its distribution.

24 keywords24 linked from6 not yet writtenWritten by AI
FunctionProbability Dist…Random VariableProbabilityStatisticsReal NumberLimitMeasure TheoryCumulative…

A cumulative distribution function (CDF) is a function that describes the probability distribution of a random variable by accumulating probability up to a given threshold. For a real-valued random variable XX, it is defined by FX(x)=P(X≤x)F_X(x)=P(X\leq x). The same definition applies to discrete, continuous, and mixed distributions, making the CDF a common language for describing uncertainty in statistics and probability theory. (online.stat.psu.edu)

Definition and fundamental properties

For every real number xx,

FX(x)=P(X≤x).F_X(x)=P(X\leq x).

Here XX is the random variable, whereas xx is a fixed threshold. The value FX(x)F_X(x) lies between zero and one; it is not the probability that XX equals xx. Every CDF satisfies four fundamental conditions:

  • Nondecreasing: if a<ba<b, then FX(a)≤FX(b)F_X(a)\leq F_X(b).
  • Right-continuous: lim⁡h↓0FX(x+h)=FX(x)\lim_{h\downarrow0}F_X(x+h)=F_X(x).
  • Lower boundary limit: lim⁡x→−∞FX(x)=0\lim_{x\to-\infty}F_X(x)=0.
  • Upper boundary limit: lim⁡x→+∞FX(x)=1\lim_{x\to+\infty}F_X(x)=1.

Conversely, any function satisfying these conditions is the CDF of some real-valued random variable. Right-continuity reflects the inclusion of the endpoint in X≤xX\leq x; a CDF need not be continuous from the left. (statlect.com)

In measure theory, a CDF uniquely determines a probability measure on the real line equipped with its Borel sigma-algebra. Thus, knowing the CDF determines probabilities not only for threshold events but for all Borel-measurable sets. Different random variables may share a CDF without being identical or defined on the same probability space. (live.ocw.mit.edu)

Recovering probabilities

For a<ba<b, subtracting accumulated probabilities gives

P(a<X≤b)=FX(b)−FX(a).P(a<X\leq b)=F_X(b)-F_X(a).

Define the left-hand limit by FX(x−)=lim⁡t↑xFX(t)F_X(x^-)=\lim_{t\uparrow x}F_X(t). Then

P(X=x)=FX(x)−FX(x−),P(X=x)=F_X(x)-F_X(x^-),

and

P(a≤X≤b)=FX(b)−FX(a−).P(a\leq X\leq b)=F_X(b)-F_X(a^-).

Consequently, the height of a jump at xx equals the probability concentrated at that point. If the CDF is continuous at xx, then P(X=x)=0P(X=x)=0. Endpoint distinctions therefore matter for distributions with point masses, although they disappear at endpoints having zero probability. (live.ocw.mit.edu)

The complementary probability,

P(X>x)=1−FX(x),P(X>x)=1-F_X(x),

is commonly called the survival function. It describes the probability of exceeding a threshold, rather than falling at or below it. This connection is particularly useful when the variable represents a lifetime or failure time. (itl.nist.gov)

Discrete and continuous distributions

For a discrete random variable with probability mass function p(t)=P(X=t)p(t)=P(X=t),

FX(x)=∑t≤xp(t).F_X(x)=\sum_{t\leq x}p(t).

For example, a Bernoulli distribution with P(X=1)=pP(X=1)=p and P(X=0)=1−pP(X=0)=1-p has

FX(x)={0,x<0,1−p,0≤x<1,1,x≥1.F_X(x)= \begin{cases} 0,&x<0,\\ 1-p,&0\leq x<1,\\ 1,&x\geq1. \end{cases}

Its CDF has jumps at zero and one. More generally, discrete distributions accumulate probability through jumps, with constant sections wherever an interval contains no possible outcomes. (statlect.com)

For an absolutely continuous distribution with probability density function ff, accumulation is expressed as an integral:

FX(x)=∫−∞xf(t) dt.F_X(x)=\int_{-\infty}^{x}f(t)\,dt.

Conversely, its derivative equals the density almost everywhere. A CDF value is a probability, whereas a density value is not itself a probability and may exceed one. Interval probabilities correspond to areas under the density curve. (online.stat.psu.edu)

For the uniform distribution on [0,1][0,1], the CDF equals zero below zero, xx between zero and one, and one above one. For the standard normal distribution, its CDF is conventionally denoted Φ\Phi, with numerical tables providing accumulated probabilities at specified thresholds. (itl.nist.gov)

A continuous CDF does not necessarily imply the existence of a density. Singular continuous distributions provide counterexamples. A mixed distribution can instead combine point masses and an absolutely continuous component; its CDF then combines jumps with continuous increases. The CDF remains defined in all these cases. (live.ocw.mit.edu)

Quantiles and random sampling

A quantile function reverses the direction of the CDF: it maps a probability level to a threshold. A standard generalized inverse is

Q(u)=inf⁡{x∈R:FX(x)≥u},0<u<1.Q(u)=\inf\{x\in\mathbb R:F_X(x)\geq u\}, \qquad 0<u<1.

Unlike an ordinary inverse, this definition works when the CDF has flat sections or jumps. For a continuous, strictly increasing CDF, it agrees with the usual inverse. Quantiles identify thresholds associated with specified cumulative probabilities, including the conventional lower median Q(0.5)Q(0.5). (www2.stat.duke.edu)

In inverse transform sampling, a uniform random variable UU on (0,1)(0,1) is transformed into Q(U)Q(U), which has CDF FXF_X. This construction works for arbitrary univariate distributions. The related probability integral transform states that FX(X)F_X(X) is uniformly distributed when FXF_X is continuous; without continuity, that conclusion does not generally hold. (www2.stat.duke.edu)

Empirical estimation and distribution testing

For observations x1,…,xnx_1,\ldots,x_n, the empirical distribution function is

F^n(x)=1n∑i=1n1{xi≤x},\widehat F_n(x)=\frac1n\sum_{i=1}^{n}\mathbf1\{x_i\leq x\},

where the indicator equals one when its condition is true and zero otherwise. It records the fraction of observations at or below each threshold. Its graph is a staircase: each distinct observation contributes a jump proportional to its frequency, so repeated values produce larger jumps. (itl.nist.gov)

CDF comparisons support statistical hypothesis testing. The Kolmogorov–Smirnov test compares an empirical CDF with a specified theoretical CDF using their maximum vertical separation. Standard one-sample critical values assume a fully specified continuous distribution; estimating its parameters from the same observations requires an adjusted calibration rather than unmodified critical values. (itl.nist.gov)