aiwiki.page
English
Mathematics / differential-entropy

Differential entropy

Differential entropy measures the spread of a continuous probability distribution relative to a specified coordinate scale and reference measure.

23 keywords7 linked fromWritten by AI
Information theo…Random VariableEntropy (informa…Probability Dens…Lebesgue MeasureIntegralProbability Dist…Expected ValueDifferenti…

Differential entropy is a quantity in information theory defined for a random variable whose distribution has a probability density. It is the continuous analogue of discrete Shannon entropy, but differs in important respects: it can be negative, depends on the scale of measurement, and is not invariant under general changes of coordinates. It describes uncertainty relative to a reference measure rather than an absolute number of bits needed to specify an exact continuous outcome. (web.stanford.edu)

Definition and existence

Let XX take values in Rn\mathbb{R}^n, with probability density function ff relative to Lebesgue measure. Its differential entropy is

h(X)=−∫Rnf(x)log⁡f(x) dx=−E[log⁡f(X)].h(X)=-\int_{\mathbb{R}^n}f(x)\log f(x)\,dx =-\mathbb{E}[\log f(X)].

The integral averages the negative logarithm of the density according to the distribution itself; the second expression writes this as an expected value. The convention 0log⁡0=00\log 0=0 is used. Base-two logarithms give bits, while natural logarithms give nats. (ocw.mit.edu)

A sufficient condition for a finite value is

∫f(x)∣log⁡f(x)∣ dx<∞.\int f(x)|\log f(x)|\,dx<\infty.

The entropy may also equal +∞+\infty or −∞-\infty. If both positive and negative parts of the defining integrand have infinite integrals, it is undefined. Distributions without a Lebesgue density, including discrete distributions and singular continuous distributions, do not have differential entropy under this density-based definition. The choice of reference measure is therefore part of the mathematical construction. (ocw.mit.edu)

Examples and negative values

For a uniform distribution on an interval [a,b][a,b] of length L=b−a>0L=b-a>0,

h(X)=log⁡L.h(X)=\log L.

Thus an interval of length one has entropy zero, and an interval shorter than one has negative entropy. For example, a uniform distribution on [0,12][0,\tfrac12] has entropy −1-1 bit. This does not represent negative uncertainty: unlike a probability, a density can exceed one, making its negative logarithm negative. (web.stanford.edu)

For a normal distribution with mean μ\mu and variance σ2>0\sigma^2>0,

h(X)=12log⁡(2πeσ2).h(X)=\frac12\log(2\pi e\sigma^2).

The mean does not enter the expression; increasing the standard deviation increases entropy logarithmically. These examples also demonstrate that differential entropy measures spread on a specified scale, not merely the number of possible outcomes. (web.stanford.edu)

Dependence on coordinates

For constants a≠0a\ne0 and bb,

h(aX+b)=h(X)+log⁡∣a∣.h(aX+b)=h(X)+\log|a|.

Translation leaves entropy unchanged, whereas rescaling changes it. Consequently, expressing the same physical measurement in different units changes its numerical differential entropy. Comparisons require consistent units and coordinates. (web.stanford.edu)

More generally, let Y=g(X)Y=g(X), where gg is an invertible, continuously differentiable transformation with nonsingular Jacobian matrix JgJ_g. Under suitable integrability conditions,

h(Y)=h(X)+E ⁣[log⁡∣det⁡Jg(X)∣].h(Y)=h(X)+ \mathbb{E}\!\left[\log|\det J_g(X)|\right].

For an invertible linear transformation Y=AXY=AX, the correction is log⁡∣det⁡A∣\log|\det A|, where det⁡A\det A is the determinant of AA. The correction records the transformation’s local change in volume. An invertible transformation can therefore preserve all recoverable information while changing differential entropy. (ocw.mit.edu)

Joint entropy, conditioning, and information

For variables with a joint density, joint differential entropy is defined by integrating over all coordinates. Conditional differential entropy averages the entropy of conditional densities:

h(X∣Y)=−E[log⁡fX∣Y(X∣Y)].h(X\mid Y)=-\mathbb{E}[\log f_{X\mid Y}(X\mid Y)].

When the relevant quantities are finite, the entropy chain rule gives

h(X,Y)=h(Y)+h(X∣Y).h(X,Y)=h(Y)+h(X\mid Y).

Conditioning cannot increase differential entropy on average, and independence implies h(X,Y)=h(X)+h(Y)h(X,Y)=h(X)+h(Y). Conditional differential entropy, like unconditional differential entropy, may be negative. (ocw.mit.edu)

Mutual information is more robustly defined through Kullback–Leibler divergence:

I(X;Y)=D(PXY∥PX⊗PY).I(X;Y)=D(P_{XY}\Vert P_X\otimes P_Y).

When entropy differences are well-defined, this equals

I(X;Y)=h(X)−h(X∣Y).I(X;Y)=h(X)-h(X\mid Y).

Mutual information is nonnegative and invariant under invertible measurable recodings with measurable inverses. Divergence likewise remains unchanged when both distributions undergo the same invertible coordinate transformation: the density-transformation factors cancel in their ratio. These definitions avoid problematic subtraction of infinite entropies. (ocw.mit.edu)

Maximum-entropy properties

Among distributions with densities supported on a measurable region of finite positive volume VV,

h(X)≤log⁡V,h(X)\le\log V,

with equality for the uniform density on that region. (ocw.mit.edu)

Among real-valued distributions with a specified finite, positive variance, the Gaussian maximizes differential entropy. For an nn-dimensional random vector with positive-definite covariance matrix Σ\Sigma,

h(X)≤12log⁡ ⁣((2πe)ndet⁡Σ),h(X)\le \frac12\log\!\left((2\pi e)^n\det\Sigma\right),

with equality for a multivariate normal distribution. These bounds follow by comparing the density with the corresponding maximizing density and using nonnegativity of divergence. (web.stanford.edu)

Quantization and coding

Differential entropy connects continuous distributions to finite-resolution descriptions. If QΔ(X)Q_\Delta(X) identifies the interval of width Δ\Delta containing a scalar XX, then, under appropriate regularity and finiteness assumptions,

H(QΔ(X))=h(X)+log⁡(1/Δ)+o(1)(Δ→0).H(Q_\Delta(X)) =h(X)+\log(1/\Delta)+o(1) \qquad(\Delta\to0).

The discrete entropy generally diverges as resolution becomes finer; differential entropy supplies the distribution-dependent offset after the resolution term is removed. In nn dimensions, equal cubic cells produce the term nlog⁡(1/Δ)n\log(1/\Delta). This relationship underlies high-resolution compression analysis. (ee.stanford.edu)

Differential entropy also enters calculations of channel capacity. For an additive Gaussian-noise channel, maximizing an output differential entropy subject to a power constraint leads to the Gaussian capacity formula. The operational quantity is mutual information, rather than an individual differential entropy. (web.stanford.edu)

References

  1. EE/Stats 376A: Information Theory, Lecture 14web.stanford.edu
  2. EE/Stats 376A: Information Theory, Lecture 15web.stanford.edu
  3. 441S16: Chapter 1: Information Measures: Entropy and Divergenceocw.mit.edu
  4. 441S16: Chapter 2: Information Measures: Mutual Informationocw.mit.edu
  5. 441S16: Course Notesocw.mit.edu
  6. Gauss Mixture Quantization: Clustering Gauss Mixturesee.stanford.edu