Conditional independence is a relation in probability theory and statistics describing variables whose dependence disappears once specified information is known. Two random variables and are conditionally independent given when their conditional joint distribution factors into their conditional marginal distributions. Written , this means that, after accounting for , observing does not change the conditional distribution of , and conversely. Unlike ordinary statistical independence, the relation explicitly depends on the information being conditioned on. (stats.ox.ac.uk)
Mathematical definition
For discrete variables, conditional independence requires
for every and every with . Equivalently,
whenever the conditioning event has positive probability. Thus, the conditional joint distribution is a product distribution within each relevant stratum of . For variables with suitable densities, the analogous factorization uses a conditional probability density:
These identities need hold only almost everywhere, rather than at arbitrarily chosen points where conditional distributions may be undefined. (stats.ox.ac.uk)
A formulation using conditional expectation avoids requiring densities. For all bounded measurable functions ,
In measure theory, conditioning on means conditioning on the sigma-algebra generated by . The same definition extends to independence conditional on a general information sigma-algebra. (stats.ox.ac.uk)
Conditioning can remove or create dependence
Marginal independence and conditional independence do not imply one another. A shared underlying variable can make two observations dependent before conditioning, even when they are independent given that variable. Conversely, conditioning on information jointly determined by two independent variables can make them dependent. These patterns correspond to common-cause and common-effect structures in graphical models. (cs.cmu.edu)
As an explicit illustration, let be independent fair binary variables and set . Conditional on , only and are possible, each with probability . Consequently,
whereas the product of the two conditional marginal probabilities is . Knowing the sum therefore destroys their independence.
Conditional independence is also stronger than zero conditional covariance or absence of conditional linear correlation. When the necessary moments exist, it implies
but this single moment identity does not generally establish independence of the full conditional distributions. (arxiv.org)
Logical properties
Conditional independence obeys four fundamental inference rules, often called the semigraphoid properties. Here may represent collections of variables:
- Symmetry: implies .
- Decomposition: implies and .
- Weak union: implies .
- Contraction: together with implies . (stats.ox.ac.uk)
A further rule, intersection, holds under suitable strict-positivity assumptions:
Without those assumptions it can fail. Weak union does not license adding arbitrary conditioning variables: its premise requires independence from the entire pair . (stats.ox.ac.uk)
Graphical representation
A probabilistic graphical model encodes conditional independence through graph structure. In a Bayesian network, a directed acyclic graph supports the factorization
where denotes the parents of node . Each variable is conditionally independent of its nondescendants given its parents. (cs.cmu.edu)
The d-separation criterion identifies further independences implied by this factorization. Conditioning on the middle node blocks a simple chain or fork . In a collider , however, conditioning on , or a descendant of , can open a previously blocked path. (cs.cmu.edu)
D-separation guarantees independence for distributions that factorize according to the graph; failure of d-separation does not guarantee dependence in every particular distribution. Special parameter choices may produce additional independences. Graphical dependence also does not by itself establish causation: causal inference requires additional assumptions about causal structure. (cs.cmu.edu)
Statistical applications and testing
In machine learning, the naive Bayes classifier assumes mutual conditional independence of features given a class variable :
Combined with Bayes’ theorem, this yields a compact classification model. The assumption concerns independence within classes, not independence across the pooled population. (cs.cmu.edu)
Conditional independence also expresses the role of a sufficient statistic: under an appropriate Bayesian interpretation with random parameter , sufficiency can be expressed as . It formalizes the idea that the statistic retains the sample’s information about the parameter. (stats.ox.ac.uk)
Empirically assessing the relation is a problem in statistical hypothesis testing. Categorical variables permit contingency-table approaches, while continuously valued conditioning variables make distribution-free testing substantially harder. General impossibility results show that useful tests require restrictions on the distributional class or other assumptions; a test of conditional covariance alone cannot generally certify conditional independence. (stat.cmu.edu)