Conditional probability is the probability assigned to an event when another event or specified information is taken as given. Written , it is read “the probability of given .” Conditioning restricts attention to outcomes consistent with the information and adjusts their probabilities accordingly. It is a fundamental tool in statistics for describing dependence and reasoning under uncertainty. (pages.stat.wisc.edu)
Definition and interpretation
Let and be events in a probability space . If , then
Here is the event that both occur. The denominator rescales the probability within , making itself have conditional probability one. Equivalently, becomes the effective sample space. For fixed , the map obeys the usual probability axioms, including nonnegativity and countable additivity. (arxiv.org)
For example, consider a fair six-sided die. Let mean “the result exceeds three” and mean “the result is even.” Within , two of the three equally likely outcomes belong to , so , whereas . This calculation illustrates restriction followed by normalization. Counting outcomes works directly only when the relevant outcomes are equally likely. (live.ocw.mit.edu)
Multiplication and total probability
Rearranging the definition gives the multiplication rule:
Repeated application gives
provided the conditioning events have positive probability. This factorization expresses a joint probability through successive conditional probabilities; it does not require independence. (arxiv.org)
The law of total probability combines conditional probabilities across alternative cases. If are mutually exclusive events covering the sample space, each with positive probability, then
Thus an unconditional probability is a weighted average of probabilities within the separate cases. The weights are their probabilities, not necessarily equal fractions. (live.ocw.mit.edu)
Bayes’ theorem and reversed conditioning
In general, and differ. Bayes’ theorem relates them:
when the quantities on the right are defined. For a partition of possible hypotheses ,
In Bayesian inference, the prior distribution represents uncertainty before incorporating evidence, the likelihood function describes the evidence under each hypothesis, and the posterior distribution represents uncertainty after conditioning. (ocw.mit.edu)
As a calculated illustration, suppose a coin is selected with equal probability from a fair coin and a coin that always produces heads. After observing one head, the probability that the selected coin is the always-heads coin is
Although heads is certain under that hypothesis, the hypothesis is not certain after heads. Reversing the conditioning without accounting for the prior probabilities would confuse two different questions. (ocw.mit.edu)
Independence and conditional independence
Events have statistical independence when
For , this is equivalent to : learning that occurred leaves the probability of unchanged. Independence is different from mutual exclusivity. Two mutually exclusive events with positive probabilities cannot be independent, because their intersection has probability zero. (pages.stat.wisc.edu)
Conditional independence is independence within a specified information state. Given an event of positive probability, it means
Neither ordinary independence nor conditional independence implies the other in general. Conditioning can therefore reveal or remove statistical dependence rather than simply preserve it. (live.ocw.mit.edu)
Conditional distributions
For discrete random variables and , their joint probability distribution determines the conditional probability mass function:
For each qualifying , these conditional probabilities sum to one. Their weighted average over recovers the distribution of . (ocw.mit.edu)
If and have a joint probability density function, a corresponding conditional density is
Conditional probabilities are obtained by integrating this density over the desired values of . For continuously distributed , the event has probability zero, so this construction is not an application of the elementary event-ratio formula. The conditional distribution also determines the conditional expectation , its expected value when that value of is given. (ocw.mit.edu)
Measure-theoretic formulation
In measure theory, information is represented by a sub-sigma-algebra . Conditional probability is defined through conditional expectation:
where equals one on and zero elsewhere. This is generally a random variable, not a single number. It is -measurable and satisfies
It is unique up to changes on sets of probability zero. This formulation accommodates conditioning on random variables without requiring their individual values to have positive probability; values on null sets are not uniquely determined by the defining identity. (arxiv.org)