A joint probability distribution is a probability distribution for two or more random variables considered together. It assigns probabilities to combinations of their values or to regions of their combined value space. Unlike separate distributions for individual variables, it describes how their outcomes occur together and therefore captures dependence between them. Joint distributions underpin multivariate statistics and probabilistic modeling. (online.stat.psu.edu)
Definition and representations
For real-valued variables and , the joint cumulative distribution function is
The comma denotes “and”: both inequalities must hold. This function characterizes the joint distribution and extends naturally to any finite number of variables. For example, three variables have the distribution function . (openlearninglibrary.mit.edu)
In the framework of measure theory, variables defined on a common probability space form a random vector. Its joint law assigns each measurable region the probability that the vector lies in . This formulation accommodates discrete, continuous, and other distributions without requiring a density. (ocw.mit.edu)
Discrete joint distributions
When and take countably many possible values, their joint probability mass function is
It satisfies
Probabilities of events are obtained by summing over all value pairs belonging to the event. With finite value sets, the masses can be displayed in a table whose rows and columns correspond to the two variables. Impossible combinations receive probability zero. (online.stat.psu.edu)
For two independent fair six-sided dice, let and denote their respective results. Each of the 36 ordered pairs has probability . Six pairs have equal coordinates, giving . The distribution of the total follows by adding probabilities along the table’s diagonals: totals with more contributing pairs are more likely. (openlearninglibrary.mit.edu)
Continuous joint distributions
If the joint law has a probability density function with respect to two-dimensional Lebesgue measure, then
The density is nonnegative and its integral over the whole plane equals one. A density value is not itself a probability and may exceed one; probabilities arise from integration over regions. Where the density is continuous, it can be recovered from the joint distribution function through partial differentiation:
Continuous individual distributions do not guarantee a joint density. For example, if is continuously distributed and , every outcome lies on the diagonal , a set of zero planar area. The joint law consequently cannot have an ordinary two-dimensional density. A joint distribution function still exists. (ocw.mit.edu)
Marginal and conditional distributions
A marginal distribution describes one variable while disregarding the other. It is obtained by summing or integrating out the unwanted coordinate:
The word “marginal” comes from writing row and column totals in the margins of probability tables. Marginalization preserves individual behavior but generally loses information about dependence. (openlearninglibrary.mit.edu)
A conditional distribution describes one variable given information about another. Using conditional probability, the discrete case gives
when . For jointly absolutely continuous variables, an analogous conditional density is
where the denominator is positive. This density formula defines conditioning through the joint law, rather than dividing by the probability of the zero-probability event . Such relationships also underlie Bayes’ theorem. (online.stat.psu.edu)
Independence and dependence
Statistical independence means that the joint law factors into the individual laws. For discrete variables,
for variables with a joint density, the corresponding density factorization holds almost everywhere. Independence is also equivalent to for all . (openlearninglibrary.mit.edu)
Marginals alone do not ordinarily determine the joint distribution. As an illustrative construction, two fair binary variables may be independent, giving four pairs probability , or always equal, giving only and probability . Both constructions have identical marginals but different joint behavior.
When the relevant moments exist, covariance summarizes part of this behavior:
Here expected values are computed from the joint law. Independence implies zero covariance, but zero covariance does not generally imply independence; correlation measures linear association rather than all forms of dependence. (ocw.mit.edu)
Higher-dimensional models and computation
For discrete variables, repeated conditioning yields
with conditionals interpreted on configurations of positive probability. This probability chain rule requires no independence assumption. Conditional independence can simplify its factors. A Bayesian network, one kind of probabilistic graphical model, represents a joint law using conditional distributions associated with each variable’s parents. (cs.cmu.edu)
Explicit tables grow rapidly: variables with possible values require entries before normalization constraints. Structured factorizations can reduce storage and support variable elimination, which answers queries by combining factors and summing out unneeded variables without first constructing the complete table. (cs.cmu.edu)