Kullback–Leibler divergence, usually abbreviated KL divergence, is a directional measure of discrepancy between two probability distributions. It expresses the expected logarithmic ratio of their probabilities, with the expectation taken under the first distribution. Central to information theory, statistics, and machine learning, it is also called relative entropy. Solomon Kullback and Richard Leibler introduced the measure in their 1951 paper “On Information and Sufficiency.” Although often described informally as a distance, it is not a mathematical metric. (www-ee.stanford.edu)
Definition
For discrete distributions and on the same finite or countable set, with probability mass functions and ,
The order matters: the weights come from , while supplies the comparison probabilities. A term with contributes zero, including when . If but , the divergence is infinite. Natural logarithms give units called nats; base-two logarithms give bits. Changing the logarithm base rescales the result by a constant. (web.stanford.edu)
For continuous distributions with densities relative to a common reference measure,
In measure theory, the general definition uses the Radon–Nikodym derivative:
provided is absolutely continuous with respect to ; otherwise it is infinite. Absolute continuity means that every event assigned zero probability by also has zero probability under . Even when this condition holds, the integral can diverge. (www-ee.stanford.edu)
Information-theoretic interpretation
KL divergence is the expected value under of a log-likelihood ratio. An observation favored more strongly by than by contributes positively; an observation favored by contributes negatively. Individual contributions can therefore be negative even though their overall expectation cannot. (web.stanford.edu)
For discrete distributions with finite relevant entropies, it relates entropy and cross-entropy:
where
In lossless data compression, this difference represents the expected excess ideal code length incurred by encoding outcomes generated by using probabilities from . Literal integer-length codes introduce rounding overhead, so the interpretation is exact for ideal lengths and appropriate asymptotic coding settings rather than every individual code. (theory.stanford.edu)
For continuous densities, an analogous entropy-difference identity requires suitable finiteness conditions. Unlike differential entropy, KL divergence is invariant under a common invertible change of coordinates: the density transformation factors cancel in the ratio. (www-ee.stanford.edu)
Mathematical properties
Non-negativity, commonly expressed as Gibbs’ inequality, states
with equality precisely when as probability measures. Nevertheless, KL divergence is generally asymmetric and does not satisfy the triangle inequality. It therefore differs fundamentally from Euclidean distance. (web.stanford.edu)
Several further properties make it useful in statistical analysis:
- Joint convexity: mixing corresponding pairs of distributions cannot increase divergence beyond the same mixture of their divergences.
- Additivity: for independent product distributions, the joint divergence equals the sum of component divergences.
- Data processing: applying the same measurable transformation or random channel to both distributions cannot increase their divergence.
These results connect KL divergence with convex optimization and formalize how aggregation or discarded information limits statistical distinguishability. (stanford.edu)
For random variables and , mutual information is a particular KL divergence:
It compares their joint distribution with the product distribution that would describe independence. (stanford.edu)
Direction and examples
Consider distributions on two outcomes, and . Direct substitution gives
The second result occurs because assigns positive probability to an outcome that excludes. This illustrates why reversing the arguments changes both the weighting and the support requirements.
For a target density and a restricted approximation , minimizing penalizes assigning too little probability wherever the target has mass. Minimizing instead averages over the approximation and strongly penalizes placing mass where the target is very small. With multimodal targets, these directions can produce broader coverage or concentration around one mode, respectively; such behavior depends on the available approximation family. (cs.columbia.edu)
Statistical and machine-learning applications
For a fixed data distribution, minimizing cross-entropy also minimizes KL divergence because its entropy term is constant. Accordingly, maximum likelihood estimation for discrete observations can be expressed as minimizing divergence from the empirical distribution to a model. Classification commonly uses this relationship to construct a loss function from predicted class probabilities. (nlp.stanford.edu)
In Bayesian inference, variational inference often approximates a posterior by minimizing . Equivalently, it maximizes the evidence lower bound:
This formulation avoids directly optimizing an objective containing the generally intractable evidence . (cs.columbia.edu)
A variational autoencoder combines an expected reconstruction term with a KL penalty between its approximate posterior and latent prior. This penalty provides regularization of the latent representation. When both distributions are suitable Gaussian distributions, the divergence can be evaluated analytically, while reconstruction expectations may require sampling. (arxiv.org)