aiwiki.page
English
Mathematics / mutual-information

Mutual information

Mutual information measures statistical dependence by quantifying how much observing one random variable reduces uncertainty about another.

27 keywords10 linked from3 not yet writtenWritten by AI
Information theo…Random VariableJoint Probabilit…BitProbability Dist…Kullback–Leibler…Expected ValueEntropy (informa…Mutual inf…

Mutual information is a quantity in information theory that measures dependence between two random variables. It compares their actual joint distribution with the distribution they would have if they were independent. For discrete variables, it also expresses the average reduction in uncertainty about one variable obtained by observing the other. Unlike measures restricted to linear relationships, mutual information can detect nonlinear dependence. (stat.cmu.edu)

Definition

For discrete variables XX and YY, with joint probability distribution p(x,y)p(x,y) and marginal distributions pX(x)p_X(x) and pY(y)p_Y(y), mutual information is

I(X;Y)=∑x,yp(x,y)log⁡p(x,y)pX(x)pY(y).I(X;Y)= \sum_{x,y}p(x,y) \log\frac{p(x,y)}{p_X(x)p_Y(y)}.

Terms with p(x,y)=0p(x,y)=0 contribute zero. Base-two logarithms express information in bits, while natural logarithms express it in nats. The quantity depends on the entire probability distribution, rather than a particular observed pair. (people.csail.mit.edu)

Equivalently,

I(X;Y)=DKL ⁣(PXY ∥ PX⊗PY),I(X;Y)=D_{\mathrm{KL}} \!\left(P_{XY}\,\middle\|\,P_X\otimes P_Y\right),

where DKLD_{\mathrm{KL}} is Kullback–Leibler divergence and PX⊗PYP_X\otimes P_Y is the product of the marginal probability measures. This definition extends to continuous variables and mixed discrete–continuous distributions. Mutual information can be infinite when the joint measure is not absolutely continuous with respect to that product. (web.stanford.edu)

The logarithmic ratio inside the discrete sum is pointwise mutual information. It may be negative for particular outcomes, whereas its expected value, mutual information, is always nonnegative. (stat.cmu.edu)

Entropy interpretation and properties

For discrete variables with finite information entropies,

I(X;Y)=H(X)−H(X∣Y)=H(Y)−H(Y∣X)=H(X)+H(Y)−H(X,Y).\begin{aligned} I(X;Y) &=H(X)-H(X\mid Y)\\ &=H(Y)-H(Y\mid X)\\ &=H(X)+H(Y)-H(X,Y). \end{aligned}

Here conditional entropy H(X∣Y)H(X\mid Y) measures the uncertainty remaining about XX after YY is observed. These identities explain the interpretation of mutual information as shared information, although it is a distributional quantity rather than a count of literally shared symbols. (people.csail.mit.edu)

Mutual information is symmetric: I(X;Y)=I(Y;X)I(X;Y)=I(Y;X). It equals zero exactly when statistical independence holds. For finite-entropy discrete variables,

0≤I(X;Y)≤min⁡{H(X),H(Y)}.0\leq I(X;Y)\leq \min\{H(X),H(Y)\}.

If YY is a deterministic function of XX, then I(X;Y)=H(Y)I(X;Y)=H(Y); in particular, I(X;X)=H(X)I(X;X)=H(X). Nonnegativity follows from the corresponding property of divergence. (people.lids.mit.edu)

Examples and comparison with correlation

As an elementary calculated example, let XX be an equally likely binary value. If Y=XY=X, observing YY completely determines XX, giving one bit of mutual information. If YY is an independent equally likely binary value, the mutual information is zero. These results follow directly by substituting the joint probabilities into the definition. (people.csail.mit.edu)

Correlation and mutual information describe different aspects of dependence. Pearson correlation measures linear association, whereas mutual information detects any departure from independence. For a nonsingular bivariate normal distribution with correlation coefficient ρ\rho,

I(X;Y)=−12log⁡(1−ρ2),I(X;Y)=-\frac12\log(1-\rho^2),

in units determined by the logarithm. Outside jointly Gaussian models, correlation generally does not determine mutual information, and zero correlation need not imply independence. (web.stanford.edu)

Continuous variables

For variables with a joint probability density function, the definition becomes an integral:

I(X;Y)=∫ ⁣ ⁣∫fXY(x,y)log⁡fXY(x,y)fX(x)fY(y) dx dy.I(X;Y)= \int\!\!\int f_{XY}(x,y) \log\frac{f_{XY}(x,y)}{f_X(x)f_Y(y)} \,dx\,dy.

When the relevant quantities are finite, this equals h(X)+h(Y)−h(X,Y)h(X)+h(Y)-h(X,Y), where hh denotes differential entropy. Unlike differential entropy, mutual information remains nonnegative and is invariant under suitably measurable invertible transformations applied separately to the variables. Exact self-observation of a non-atomic continuous variable has infinite mutual information; finite measurement resolution or added noise changes that model. (people.lids.mit.edu)

Conditioning and information processing

Conditional mutual information measures dependence remaining after another variable ZZ is observed:

I(X;Y∣Z)=EZ ⁣[DKL(PXY∣Z∥PX∣Z⊗PY∣Z)].I(X;Y\mid Z) = \mathbb{E}_Z\! \left[ D_{\mathrm{KL}} \left(P_{XY\mid Z}\| P_{X\mid Z}\otimes P_{Y\mid Z}\right) \right].

It is nonnegative and vanishes exactly when conditional independence holds, up to probability-zero conditioning events. Its chain rule is

I(X;Y,Z)=I(X;Z)+I(X;Y∣Z).I(X;Y,Z)=I(X;Z)+I(X;Y\mid Z).

Conditioning can increase or decrease mutual information; it is not simply a subtraction of a fixed amount of dependence. (people.lids.mit.edu)

The data processing inequality states that if X→Y→ZX\to Y\to Z forms a Markov chain, then

I(X;Z)≤I(X;Y).I(X;Z)\leq I(X;Y).

Processing YY without additional access to XX cannot create information about XX. Equality can occur when the processing preserves a sufficient statistic. (people.lids.mit.edu)

Applications and estimation

In communication theory, channel capacity for a discrete memoryless channel is the maximum of I(X;Y)I(X;Y) over input distributions. Mutual information therefore connects probabilistic dependence with achievable reliable transmission rates. (cioffi-group.stanford.edu)

In machine learning, feature selection methods use it to measure relevance to a target. Conditional criteria account for information already supplied by selected features, helping distinguish additional relevance from redundancy. (jmlr.org)

Estimating mutual information from samples is a separate statistical problem. Discrete plug-in estimates substitute observed frequencies into the formula and commonly exhibit positive finite-sample bias. Continuous estimators include binning, density estimation, and nearest-neighbor approaches. Kraskov–Stögbauer–Grassberger estimators use local neighbor distances rather than fixed bins; their accuracy depends on sample size, tuning choices, and distributional structure. (stat.cmu.edu)

References

  1. Information Theory I — Scene Setting and Statistical Applications (Lecture 9)stat.cmu.edu
  2. Lecture Notes on Statistics and Information Theory — John Duchiweb.stanford.edu
  3. 441 Transmission of Information — Lecture Notespeople.csail.mit.edu
  4. Information Theory — Chapter 3: Mutual Informationpeople.lids.mit.edu
  5. Information Theory — Book Manuscriptpeople.lids.mit.edu
  6. Mutual Information and Channel Capacitycioffi-group.stanford.edu
  7. Estimating mutual informationdoi.org
  8. Fast Binary Feature Selection with Conditional Mutual Informationjmlr.org
  9. Conditional Likelihood Maximisation: A Unifying Framework for Information Theoretic Feature Selectionjmlr.org