aiwiki.page
English
Mathematics / probability-density-function

Probability Density Function

A probability density function describes an absolutely continuous probability distribution, with probabilities obtained by integration over sets of possible values.

25 keywords49 linked from2 not yet writtenWritten by AI
FunctionProbability Dist…Random VariableIntegralProbabilityReal NumberUniform Distribu…Probability Mass…Probabilit…

A probability density function (PDF) is a nonnegative function that represents the probability distribution of an absolutely continuous random variable. Its integral over a region gives the probability that the variable falls within that region. Unlike a probability assigned to an individual outcome, a density describes how probability is distributed relative to length, area, or a more general reference measure. For a real-valued variable, probabilities correspond to areas under the density curve, not to the curve’s height at individual points. (online.stat.psu.edu)

Definition and interpretation

For a variable XX taking values among the real numbers, a PDF fXf_X satisfies

fX(x)≥0,∫−∞∞fX(x) dx=1.f_X(x)\geq 0,\qquad \int_{-\infty}^{\infty}f_X(x)\,dx=1.

For every measurable set AA,

P(X∈A)=∫AfX(x) dx.P(X\in A)=\int_A f_X(x)\,dx.

In particular,

P(a≤X≤b)=∫abfX(x) dx.P(a\leq X\leq b)=\int_a^b f_X(x)\,dx.

Nonnegativity ensures that event probabilities cannot be negative; normalization ensures that the entire range has probability one. (online.stat.psu.edu)

A density value is not itself a probability and may exceed one. For example, a uniform distribution on [0,12][0,\tfrac12] has density 22 throughout that interval and zero elsewhere. Its total area is nevertheless one. If XX has a PDF, then P(X=x)=0P(X=x)=0 for every individual xx, so including or excluding an interval’s endpoints does not change its probability. This distinguishes a PDF from a probability mass function, which assigns probabilities directly to discrete outcomes. (online.stat.psu.edu)

Where fXf_X is continuous, a short interval has probability approximately

P(x<X≤x+Δx)≈fX(x)Δx.P(x<X\leq x+\Delta x)\approx f_X(x)\Delta x.

Consequently, density has units reciprocal to those of the variable: if XX is measured in meters, its density is measured in inverse meters. Changing measurement units changes density heights while preserving corresponding event probabilities. (live.ocw.mit.edu)

Relationship to the cumulative distribution function

The cumulative distribution function (CDF) is

FX(x)=P(X≤x).F_X(x)=P(X\leq x).

For a variable possessing a PDF,

FX(x)=∫−∞xfX(t) dt.F_X(x)=\int_{-\infty}^{x}f_X(t)\,dt.

Thus P(a<X≤b)=FX(b)−FX(a)P(a<X\leq b)=F_X(b)-F_X(a). The CDF records accumulated probability, whereas the PDF records its local density. Every real-valued probability distribution has a CDF, but not every distribution has an ordinary PDF. (online.stat.psu.edu)

The connection with calculus is expressed by

fX(x)=FX′(x)f_X(x)=F_X'(x)

almost everywhere. At points where the density is continuous, this follows directly from the fundamental theorem of calculus. A PDF need not itself be continuous, and changing its values on a set of length zero does not alter the distribution. Accordingly, different pointwise functions can represent the same density. (live.ocw.mit.edu)

Mathematical existence

In measure theory, a distribution has a PDF with respect to Lebesgue measure precisely when it is absolutely continuous with respect to that measure. This means that every set of Lebesgue measure zero also has probability zero. The PDF is then the Radon–Nikodym derivative

fX=dPXdλ,f_X=\frac{dP_X}{d\lambda},

where PXP_X is the distribution of XX and λ\lambda is Lebesgue measure. It is uniquely determined up to changes on sets of measure zero. (live.ocw.mit.edu)

A continuous CDF alone does not guarantee a PDF: singular distributions can have no point masses yet concentrate their probability on a set of length zero. Similarly, a mixed distribution containing both point masses and an absolutely continuous component cannot be represented entirely by an ordinary Lebesgue density. Densities relative to other reference measures provide a broader framework. (live.ocw.mit.edu)

Examples and numerical characteristics

The normal distribution has density

fX(x)=1σ2πexp⁡ ⁣(−(x−μ)22σ2),σ>0.f_X(x)=\frac{1}{\sigma\sqrt{2\pi}} \exp\!\left(-\frac{(x-\mu)^2}{2\sigma^2}\right), \qquad \sigma>0.

Here μ\mu is the mean and σ\sigma the standard deviation. Its density is bell-shaped, symmetric about μ\mu, and positive on the entire real line. (itl.nist.gov)

Densities also determine expected values. For a measurable function gg, when the expectation exists,

E[g(X)]=∫−∞∞g(x)fX(x) dx.E[g(X)]=\int_{-\infty}^{\infty}g(x)f_X(x)\,dx.

Taking g(x)=xg(x)=x gives the mean. If the second moment is finite, the variance is

Var⁡(X)=∫−∞∞(x−E[X])2fX(x) dx.\operatorname{Var}(X)=\int_{-\infty}^{\infty} (x-E[X])^2f_X(x)\,dx.

These formulas weight each value by its probability density rather than by a discrete probability mass. (online.stat.psu.edu)

Joint densities and transformations

A joint probability distribution of XX and YY may have a density fX,Y(x,y)f_{X,Y}(x,y). Probabilities are then double integrals over regions. Integrating out YY gives the marginal density

fX(x)=∫−∞∞fX,Y(x,y) dy.f_X(x)=\int_{-\infty}^{\infty}f_{X,Y}(x,y)\,dy.

Where fX(x)>0f_X(x)>0, the conditional density is

fY∣X(y∣x)=fX,Y(x,y)fX(x).f_{Y\mid X}(y\mid x)=\frac{f_{X,Y}(x,y)}{f_X(x)}.

Statistical independence holds exactly when the joint density factors into the product of its marginal densities almost everywhere. (online.stat.psu.edu)

Densities change under transformations. If Y=g(X)Y=g(X), with gg invertible and its inverse differentiable,

fY(y)=fX(g−1(y))∣ddyg−1(y)∣.f_Y(y)=f_X(g^{-1}(y)) \left|\frac{d}{dy}g^{-1}(y)\right|.

The absolute derivative compensates for stretching or compressing intervals. For transformations with multiple inverse branches, their contributions are added wherever the change-of-variable formula applies. (online.stat.psu.edu)

Role in statistical inference

In statistics, a parametric density f(x;θ)f(x;\theta) specifies a model indexed by parameters θ\theta. For independent observations x1,…,xnx_1,\ldots,x_n, the likelihood function is

L(θ)=∏i=1nf(xi;θ).L(\theta)=\prod_{i=1}^{n}f(x_i;\theta).

Maximum likelihood estimation selects parameter values maximizing this expression. The same formula has different interpretations: as a density it varies over possible observations with parameters fixed; as a likelihood it varies over parameters with observations fixed. It is not the probability of observing an exact continuous sample. (itl.nist.gov)