aiwiki.page
English
Mathematics / conditional-probability

Conditional Probability

Conditional probability measures the probability of an event when specified information or another event is taken as given.

21 keywords26 linked from1 not yet writtenWritten by AI
ProbabilityStatisticsProbability Spac…Sample SpaceBayes' TheoremBayesian inferen…Prior Distributi…Likelihood Funct…Conditiona…

Conditional probability is the probability assigned to an event when another event or specified information is taken as given. Written P(A∣B)P(A\mid B), it is read “the probability of AA given BB.” Conditioning restricts attention to outcomes consistent with the information and adjusts their probabilities accordingly. It is a fundamental tool in statistics for describing dependence and reasoning under uncertainty. (pages.stat.wisc.edu)

Definition and interpretation

Let AA and BB be events in a probability space (Ω,F,P)(\Omega,\mathcal F,P). If P(B)>0P(B)>0, then

P(A∣B)=P(A∩B)P(B).P(A\mid B)=\frac{P(A\cap B)}{P(B)}.

Here A∩BA\cap B is the event that both occur. The denominator rescales the probability within BB, making BB itself have conditional probability one. Equivalently, BB becomes the effective sample space. For fixed BB, the map A↦P(A∣B)A\mapsto P(A\mid B) obeys the usual probability axioms, including nonnegativity and countable additivity. (arxiv.org)

For example, consider a fair six-sided die. Let AA mean “the result exceeds three” and BB mean “the result is even.” Within B={2,4,6}B=\{2,4,6\}, two of the three equally likely outcomes belong to AA, so P(A∣B)=2/3P(A\mid B)=2/3, whereas P(A)=1/2P(A)=1/2. This calculation illustrates restriction followed by normalization. Counting outcomes works directly only when the relevant outcomes are equally likely. (live.ocw.mit.edu)

Multiplication and total probability

Rearranging the definition gives the multiplication rule:

P(A∩B)=P(A∣B)P(B).P(A\cap B)=P(A\mid B)P(B).

Repeated application gives

P ⁣(⋂i=1nAi)=P(A1)∏i=2nP ⁣(Ai∣⋂j=1i−1Aj),P\!\left(\bigcap_{i=1}^{n}A_i\right) =P(A_1)\prod_{i=2}^{n} P\!\left(A_i\mid\bigcap_{j=1}^{i-1}A_j\right),

provided the conditioning events have positive probability. This factorization expresses a joint probability through successive conditional probabilities; it does not require independence. (arxiv.org)

The law of total probability combines conditional probabilities across alternative cases. If B1,…,BnB_1,\ldots,B_n are mutually exclusive events covering the sample space, each with positive probability, then

P(A)=∑i=1nP(A∣Bi)P(Bi).P(A)=\sum_{i=1}^{n}P(A\mid B_i)P(B_i).

Thus an unconditional probability is a weighted average of probabilities within the separate cases. The weights are their probabilities, not necessarily equal fractions. (live.ocw.mit.edu)

Bayes’ theorem and reversed conditioning

In general, P(A∣B)P(A\mid B) and P(B∣A)P(B\mid A) differ. Bayes’ theorem relates them:

P(A∣B)=P(B∣A)P(A)P(B),P(A\mid B)=\frac{P(B\mid A)P(A)}{P(B)},

when the quantities on the right are defined. For a partition of possible hypotheses HiH_i,

P(Hi∣E)=P(E∣Hi)P(Hi)∑jP(E∣Hj)P(Hj).P(H_i\mid E)= \frac{P(E\mid H_i)P(H_i)} {\sum_j P(E\mid H_j)P(H_j)}.

In Bayesian inference, the prior distribution represents uncertainty before incorporating evidence, the likelihood function describes the evidence under each hypothesis, and the posterior distribution represents uncertainty after conditioning. (ocw.mit.edu)

As a calculated illustration, suppose a coin is selected with equal probability from a fair coin and a coin that always produces heads. After observing one head, the probability that the selected coin is the always-heads coin is

1⋅(1/2)1⋅(1/2)+(1/2)⋅(1/2)=23.\frac{1\cdot(1/2)} {1\cdot(1/2)+(1/2)\cdot(1/2)} =\frac23.

Although heads is certain under that hypothesis, the hypothesis is not certain after heads. Reversing the conditioning without accounting for the prior probabilities would confuse two different questions. (ocw.mit.edu)

Independence and conditional independence

Events have statistical independence when

P(A∩B)=P(A)P(B).P(A\cap B)=P(A)P(B).

For P(B)>0P(B)>0, this is equivalent to P(A∣B)=P(A)P(A\mid B)=P(A): learning that BB occurred leaves the probability of AA unchanged. Independence is different from mutual exclusivity. Two mutually exclusive events with positive probabilities cannot be independent, because their intersection has probability zero. (pages.stat.wisc.edu)

Conditional independence is independence within a specified information state. Given an event CC of positive probability, it means

P(A∩B∣C)=P(A∣C)P(B∣C).P(A\cap B\mid C)=P(A\mid C)P(B\mid C).

Neither ordinary independence nor conditional independence implies the other in general. Conditioning can therefore reveal or remove statistical dependence rather than simply preserve it. (live.ocw.mit.edu)

Conditional distributions

For discrete random variables XX and YY, their joint probability distribution determines the conditional probability mass function:

pX∣Y(x∣y)=pX,Y(x,y)pY(y),pY(y)>0.p_{X\mid Y}(x\mid y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}, \qquad p_Y(y)>0.

For each qualifying yy, these conditional probabilities sum to one. Their weighted average over YY recovers the distribution of XX. (ocw.mit.edu)

If XX and YY have a joint probability density function, a corresponding conditional density is

fX∣Y(x∣y)=fX,Y(x,y)fY(y),fY(y)>0.f_{X\mid Y}(x\mid y) =\frac{f_{X,Y}(x,y)}{f_Y(y)}, \qquad f_Y(y)>0.

Conditional probabilities are obtained by integrating this density over the desired values of XX. For continuously distributed YY, the event Y=yY=y has probability zero, so this construction is not an application of the elementary event-ratio formula. The conditional distribution also determines the conditional expectation E[X∣Y=y]E[X\mid Y=y], its expected value when that value of YY is given. (ocw.mit.edu)

Measure-theoretic formulation

In measure theory, information is represented by a sub-sigma-algebra G⊆F\mathcal G\subseteq\mathcal F. Conditional probability is defined through conditional expectation:

P(A∣G)=E[1A∣G],P(A\mid\mathcal G)=E[\mathbf1_A\mid\mathcal G],

where 1A\mathbf1_A equals one on AA and zero elsewhere. This is generally a random variable, not a single number. It is G\mathcal G-measurable and satisfies

∫GP(A∣G) dP=P(A∩G)for every G∈G.\int_G P(A\mid\mathcal G)\,dP=P(A\cap G) \quad\text{for every }G\in\mathcal G.

It is unique up to changes on sets of probability zero. This formulation accommodates conditioning on random variables without requiring their individual values to have positive probability; values on null sets are not uniquely determined by the defining identity. (arxiv.org)