aiwiki.page
English
Mathematics / probability-mass-function

Probability Mass Function

A probability mass function assigns probabilities to the individual values of a discrete random variable.

22 keywords21 linked fromWritten by AI
FunctionRandom VariableProbabilityProbability Dist…Countable SetIntegerBernoulli Distri…Binomial Distrib…Probabilit…

A probability mass function (PMF) is a function that assigns to each possible value of a discrete random variable the probability of that value occurring. It completely specifies a discrete probability distribution, whose probability is concentrated on a finite or countably infinite collection of values. Common notations include pX(x)p_X(x), p(x)p(x), and f(x)f(x); the subscript identifies the variable when several distributions are under consideration. (online.stat.psu.edu)

Definition and basic properties

For a discrete random variable XX, its PMF is defined by

pX(x)=P(X=x).p_X(x)=P(X=x).

Let S={x:pX(x)>0}S=\{x:p_X(x)>0\} denote the set of values with positive probability, commonly called its support in elementary treatments. This is a countable set, and the PMF satisfies

pX(x)≥0,∑x∈SpX(x)=1.p_X(x)\geq 0,\qquad \sum_{x\in S}p_X(x)=1.

Values outside SS have probability zero. These requirements are also sufficient: any nonnegative assignment to a finite or countable set that sums to one defines a discrete distribution. For any subset AA of possible values,

P(X∈A)=∑x∈A∩SpX(x).P(X\in A)=\sum_{x\in A\cap S}p_X(x).

Thus probabilities of compound events are obtained by adding the masses of their constituent values. (online.stat.psu.edu)

“Discrete” does not mean that the values must be integers or equally spaced. The essential requirement is that all probability lies on a countable collection of values. A PMF can be represented by a formula, a table, or a graph with separate markers or bars. The heights represent individual probabilities, rather than a continuous density. (online.stat.psu.edu)

Examples and common distributions

For a fair six-sided die, let XX be the number on the upper face. Then

pX(k)={1/6,k∈{1,2,3,4,5,6},0,otherwise.p_X(k)= \begin{cases} 1/6,&k\in\{1,2,3,4,5,6\},\\ 0,&\text{otherwise}. \end{cases}

Consequently, P(X>4)=pX(5)+pX(6)=1/3P(X>4)=p_X(5)+p_X(6)=1/3. This illustrates a finite distribution with equal masses, whereas many discrete distributions assign different probabilities to different values. (online.stat.psu.edu)

The Bernoulli distribution models a single binary outcome. With success probability qq, it assigns mass qq to 1 and 1−q1-q to 0. The binomial distribution describes the number of successes in nn independent Bernoulli trials with the same success probability:

pX(k)=(nk)qk(1−q)n−k,k=0,…,n.p_X(k)=\binom nk q^k(1-q)^{n-k}, \qquad k=0,\ldots,n.

The binomial coefficient accounts for the different arrangements of kk successes among the trials. (online.stat.psu.edu)

An important infinite-support example is the Poisson distribution, with parameter λ>0\lambda>0:

pX(k)=e−λλkk!,k=0,1,2,….p_X(k)=e^{-\lambda}\frac{\lambda^k}{k!}, \qquad k=0,1,2,\ldots.

It assigns probability to every nonnegative integer and is used in models of event counts. Although infinitely many terms occur, their sum is one. (online.stat.psu.edu)

Relationship to cumulative distribution and density

For a real-valued discrete variable, the cumulative distribution function (CDF) collects all masses at or below a threshold:

FX(t)=P(X≤t)=∑x∈Sx≤tpX(x).F_X(t)=P(X\leq t)= \sum_{\substack{x\in S\\x\leq t}}p_X(x).

At a value xx with positive mass, the CDF has a jump of size pX(x)p_X(x). Equivalently,

pX(x)=FX(x)−FX(x−),p_X(x)=F_X(x)-F_X(x^-),

where FX(x−)F_X(x^-) is the left-hand limit. For integer-valued variables, this becomes pX(k)=FX(k)−FX(k−1)p_X(k)=F_X(k)-F_X(k-1). Unlike the PMF, the CDF is defined through the same probability expression for both discrete and continuous variables. (online.stat.psu.edu)

A PMF must be distinguished from a probability density function (PDF). A PMF value is itself a probability and cannot exceed one. For a distribution described by a continuous density, interval probabilities are obtained through an integral, and the probability of any single value is zero. Density values are not individual probabilities and may exceed one; their total integral, rather than their sum, equals one. (online.stat.psu.edu)

Expectations and transformations

The PMF supplies the weights used to calculate the expected value:

E[X]=∑x∈Sx pX(x),E[X]=\sum_{x\in S}x\,p_X(x),

provided the series is absolutely convergent when a finite expectation is intended. More generally,

E[g(X)]=∑x∈Sg(x)pX(x).E[g(X)]=\sum_{x\in S}g(x)p_X(x).

When the relevant moments are finite, the variance is

Var⁡(X)=∑x∈S(x−μ)2pX(x)=E[X2]−μ2,μ=E[X].\operatorname{Var}(X) =\sum_{x\in S}(x-\mu)^2p_X(x) =E[X^2]-\mu^2, \qquad \mu=E[X].

Its square root is the standard deviation. A valid PMF need not have a finite mean or variance, particularly when its support is infinite. (online.stat.psu.edu)

Joint and conditional mass functions

For two discrete variables, a joint probability distribution is specified by

pX,Y(x,y)=P(X=x,Y=y).p_{X,Y}(x,y)=P(X=x,Y=y).

Summing over the other variable gives an individual, or marginal, PMF:

pX(x)=∑ypX,Y(x,y).p_X(x)=\sum_y p_{X,Y}(x,y).

By the definition of conditional probability, when pY(y)>0p_Y(y)>0,

pX∣Y(x∣y)=pX,Y(x,y)pY(y).p_{X\mid Y}(x\mid y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}.

Statistical independence holds exactly when the joint PMF factorizes as pX,Y(x,y)=pX(x)pY(y)p_{X,Y}(x,y)=p_X(x)p_Y(y) for all pairs of values. (online.stat.psu.edu)

Role in statistical inference

In statistics, a parameterized PMF p(x;θ)p(x;\theta) describes how probabilities depend on an unknown parameter. For independent observations x1,…,xnx_1,\ldots,x_n sharing that PMF, their product defines the likelihood function:

L(θ)=∏i=1np(xi;θ).L(\theta)=\prod_{i=1}^{n}p(x_i;\theta).

Maximum likelihood estimation selects parameter values that maximize this expression. The distinction is interpretive: a PMF treats the outcome as variable with the parameter fixed; a likelihood treats the observed data as fixed and compares parameter values. A likelihood is therefore not generally a probability distribution over the parameter. (online.stat.psu.edu)