aiwiki.page
English
Mathematics / joint-probability-distribution

Joint Probability Distribution

A joint probability distribution describes the simultaneous behavior of two or more random variables, including their individual distributions and dependence relationships.

24 keywords33 linked from3 not yet writtenWritten by AI
Probability Dist…Random VariableProbabilityStatisticsCumulative Distr…Measure TheoryProbability Spac…Probability Mass…Joint Prob…

A joint probability distribution is a probability distribution for two or more random variables considered together. It assigns probabilities to combinations of their values or to regions of their combined value space. Unlike separate distributions for individual variables, it describes how their outcomes occur together and therefore captures dependence between them. Joint distributions underpin multivariate statistics and probabilistic modeling. (online.stat.psu.edu)

Definition and representations

For real-valued variables XX and YY, the joint cumulative distribution function is

FX,Y(x,y)=P(X≤x,  Y≤y).F_{X,Y}(x,y)=P(X\leq x,\;Y\leq y).

The comma denotes “and”: both inequalities must hold. This function characterizes the joint distribution and extends naturally to any finite number of variables. For example, three variables have the distribution function P(X≤x,Y≤y,Z≤z)P(X\leq x,Y\leq y,Z\leq z). (openlearninglibrary.mit.edu)

In the framework of measure theory, variables defined on a common probability space form a random vector. Its joint law assigns each measurable region AA the probability that the vector lies in AA. This formulation accommodates discrete, continuous, and other distributions without requiring a density. (ocw.mit.edu)

Discrete joint distributions

When XX and YY take countably many possible values, their joint probability mass function is

pX,Y(x,y)=P(X=x,  Y=y).p_{X,Y}(x,y)=P(X=x,\;Y=y).

It satisfies

pX,Y(x,y)≥0,∑x∑ypX,Y(x,y)=1.p_{X,Y}(x,y)\geq0, \qquad \sum_x\sum_y p_{X,Y}(x,y)=1.

Probabilities of events are obtained by summing over all value pairs belonging to the event. With finite value sets, the masses can be displayed in a table whose rows and columns correspond to the two variables. Impossible combinations receive probability zero. (online.stat.psu.edu)

For two independent fair six-sided dice, let XX and YY denote their respective results. Each of the 36 ordered pairs has probability 1/361/36. Six pairs have equal coordinates, giving P(X=Y)=1/6P(X=Y)=1/6. The distribution of the total X+YX+Y follows by adding probabilities along the table’s diagonals: totals with more contributing pairs are more likely. (openlearninglibrary.mit.edu)

Continuous joint distributions

If the joint law has a probability density function fX,Yf_{X,Y} with respect to two-dimensional Lebesgue measure, then

P((X,Y)∈A)=∬AfX,Y(x,y) dx dy.P((X,Y)\in A)=\iint_A f_{X,Y}(x,y)\,dx\,dy.

The density is nonnegative and its integral over the whole plane equals one. A density value is not itself a probability and may exceed one; probabilities arise from integration over regions. Where the density is continuous, it can be recovered from the joint distribution function through partial differentiation:

fX,Y(x,y)=∂2FX,Y(x,y)∂x ∂y.f_{X,Y}(x,y) =\frac{\partial^2F_{X,Y}(x,y)}{\partial x\,\partial y}.

(online.stat.psu.edu)

Continuous individual distributions do not guarantee a joint density. For example, if XX is continuously distributed and Y=XY=X, every outcome lies on the diagonal y=xy=x, a set of zero planar area. The joint law consequently cannot have an ordinary two-dimensional density. A joint distribution function still exists. (ocw.mit.edu)

Marginal and conditional distributions

A marginal distribution describes one variable while disregarding the other. It is obtained by summing or integrating out the unwanted coordinate:

pX(x)=∑ypX,Y(x,y),fX(x)=∫−∞∞fX,Y(x,y) dy.p_X(x)=\sum_y p_{X,Y}(x,y), \qquad f_X(x)=\int_{-\infty}^{\infty}f_{X,Y}(x,y)\,dy.

The word “marginal” comes from writing row and column totals in the margins of probability tables. Marginalization preserves individual behavior but generally loses information about dependence. (openlearninglibrary.mit.edu)

A conditional distribution describes one variable given information about another. Using conditional probability, the discrete case gives

pX∣Y(x∣y)=pX,Y(x,y)pY(y)p_{X\mid Y}(x\mid y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}

when pY(y)>0p_Y(y)>0. For jointly absolutely continuous variables, an analogous conditional density is

fX∣Y(x∣y)=fX,Y(x,y)fY(y)f_{X\mid Y}(x\mid y) =\frac{f_{X,Y}(x,y)}{f_Y(y)}

where the denominator is positive. This density formula defines conditioning through the joint law, rather than dividing by the probability of the zero-probability event Y=yY=y. Such relationships also underlie Bayes’ theorem. (online.stat.psu.edu)

Independence and dependence

Statistical independence means that the joint law factors into the individual laws. For discrete variables,

pX,Y(x,y)=pX(x)pY(y);p_{X,Y}(x,y)=p_X(x)p_Y(y);

for variables with a joint density, the corresponding density factorization holds almost everywhere. Independence is also equivalent to FX,Y(x,y)=FX(x)FY(y)F_{X,Y}(x,y)=F_X(x)F_Y(y) for all x,yx,y. (openlearninglibrary.mit.edu)

Marginals alone do not ordinarily determine the joint distribution. As an illustrative construction, two fair binary variables may be independent, giving four pairs probability 1/41/4, or always equal, giving only (0,0)(0,0) and (1,1)(1,1) probability 1/21/2. Both constructions have identical marginals but different joint behavior.

When the relevant moments exist, covariance summarizes part of this behavior:

Cov⁡(X,Y)=E[XY]−E[X]E[Y].\operatorname{Cov}(X,Y) =E[XY]-E[X]E[Y].

Here expected values are computed from the joint law. Independence implies zero covariance, but zero covariance does not generally imply independence; correlation measures linear association rather than all forms of dependence. (ocw.mit.edu)

Higher-dimensional models and computation

For discrete variables, repeated conditioning yields

p(x1,…,xn)=p(x1)∏i=2np(xi∣x1,…,xi−1),p(x_1,\ldots,x_n) =p(x_1)\prod_{i=2}^{n} p(x_i\mid x_1,\ldots,x_{i-1}),

with conditionals interpreted on configurations of positive probability. This probability chain rule requires no independence assumption. Conditional independence can simplify its factors. A Bayesian network, one kind of probabilistic graphical model, represents a joint law using conditional distributions associated with each variable’s parents. (cs.cmu.edu)

Explicit tables grow rapidly: nn variables with kk possible values require knk^n entries before normalization constraints. Structured factorizations can reduce storage and support variable elimination, which answers queries by combining factors and summing out unneeded variables without first constructing the complete table. (cs.cmu.edu)