Information entropy is a numerical measure of the uncertainty associated with a random variable. In information theory, it represents the average information gained when an outcome is observed, relative to a specified probability distribution. Claude Shannon introduced its foundational formulation in his 1948 paper A Mathematical Theory of Communication. Entropy concerns the statistical predictability of outcomes, not their meaning, truth, or practical importance. (cs.yale.edu)
Definition and units
For a discrete random variable , with possible outcomes and probabilities , entropy is
The convention assigns zero contribution to impossible outcomes. The logarithm’s base determines the unit: base two gives bits, base gives nats, and base ten gives hartleys. Changing the base rescales the quantity without changing comparisons between distributions. (web.mit.edu)
An outcome’s self-information, or surprisal, is . Entropy is therefore its expected value:
Rare outcomes convey more information when observed than common ones. This does not mean that a rare outcome dominates entropy: each surprisal is weighted by its probability. Entropy belongs to the distribution as a whole, whereas surprisal belongs to an individual outcome under that distribution. (web.mit.edu)
For a binary variable with a Bernoulli distribution,
A fair coin has entropy one bit; a coin that always produces the same result has entropy zero. A coin with head probability has approximately bits of entropy per toss. These values describe average uncertainty, not the information of every individual toss. (web.mit.edu)
Mathematical properties
For a finite alphabet of outcomes,
The lower bound is attained when one outcome has probability one. The upper bound is attained by the uniform distribution. Thus, increasing the number of available outcomes does not necessarily increase entropy: their probabilities also matter. Countably infinite distributions may have infinite entropy. (cs.yale.edu)
Entropy is a concave function of the probability distribution. Consequently, mixing distributions cannot produce less entropy than their weighted average entropies. It is also unchanged by a one-to-one relabeling of outcomes. A deterministic transformation can only preserve or reduce discrete entropy, because merging outcomes discards distinctions. These properties distinguish entropy from numerical dispersion measures such as variance, which depend on the values assigned to outcomes. (arxiv.org)
Shannon motivated the logarithmic formula through requirements including continuity, increasing uncertainty for increasing numbers of equally likely choices, and consistency when a choice is decomposed into successive stages. Under these requirements, the entropy formula is determined up to a positive multiplicative constant. (cs.yale.edu)
Joint entropy, conditioning, and shared information
Joint entropy measures uncertainty about a pair of variables. Conditional entropy averages the uncertainty remaining about after observing :
The chain rule states
For independent variables, joint entropy equals the sum of their individual entropies. More generally, . (ocw.mit.edu)
Mutual information quantifies the reduction in uncertainty supplied by another variable:
It is symmetric, nonnegative, and zero exactly when the variables are independent. Conditioning cannot increase discrete entropy on average, although a particular observed value of may yield a conditional distribution with greater entropy than the original distribution. These statements concern statistical dependence, not necessarily causal relationships. (ocw.mit.edu)
Compression and sequences
Entropy has an operational interpretation in lossless data compression. For a finite discrete source, any binary uniquely decodable code has expected codeword length at least . A suitable prefix code achieves an expected length satisfying
Coding independent source symbols in blocks makes the overhead per symbol arbitrarily small. Huffman coding constructs an optimal prefix code for a specified finite distribution. These are average-length bounds, not guarantees that every message becomes shorter. (people.csail.mit.edu)
For a stochastic process, successive symbols may be dependent. Its entropy rate, when the limit exists, is
For independent, identically distributed symbols, this equals the single-symbol entropy. Dependencies can lower the rate by making future symbols predictable from preceding ones. This distinction matters for text and other structured sequences, where symbol frequencies alone do not capture all available compression. (people.csail.mit.edu)
Statistical and machine-learning applications
Cross-entropy distinguishes the true distribution from an assumed distribution :
The additional term is the Kullback–Leibler divergence. It measures the mismatch penalty in expected logarithmic prediction loss; it vanishes when the distributions coincide. (ocw.mit.edu)
In machine learning, decision-tree learning can use entropy to measure uncertainty in class labels. Information gain compares a parent node’s entropy with the weighted average entropy after a split. In language models, average negative log-probability provides a cross-entropy-based evaluation measure; perplexity is its exponential, using the corresponding logarithm base. Cross-entropy is also used as a loss function for probabilistic predictions. (cs.cmu.edu)
Continuous variables and physical entropy
For a continuous variable with density , differential entropy is
Unlike discrete entropy, it can be negative and depends on the coordinate scale: for nonzero . It is therefore not directly interchangeable with discrete entropy or a finite-resolution coding requirement. (people.csail.mit.edu)
Information entropy shares its mathematical form with physical entropy in statistical mechanics, where probabilities describe physical microstates and Boltzmann’s constant supplies physical units. Applying this connection requires a physical model; statistical unpredictability alone does not specify thermodynamic entropy or heat transfer. (cs.yale.edu)