aiwiki.page
English
Mathematics / cross-entropy

Cross-entropy

Cross-entropy measures the expected logarithmic loss incurred when one probability distribution is used to describe outcomes generated by another.

27 keywords32 linked from1 not yet writtenWritten by AI
Information theo…Machine LearningLoss functionProbability Dist…Expected ValueBitSelf-informationEntropy (informa…Cross-entr…

Cross-entropy is a quantity in information theory that measures the average negative logarithm of probabilities assigned by a model distribution to outcomes generated by a reference distribution. It connects probabilistic prediction with information coding: predictions assigning low probability to frequently occurring outcomes incur a larger cost. In machine learning, cross-entropy is widely used as a loss function for fitting probabilistic models and evaluating their predictions. (deeplearningbook.org)

Mathematical definition

For two discrete probability distributions, PP and QQ, over the same outcome space X\mathcal X, cross-entropy is

H(P,Q)=−∑x∈XP(x)log⁡Q(x)=EX∼P[−log⁡Q(X)].H(P,Q)=-\sum_{x\in\mathcal X}P(x)\log Q(x) =\mathbb E_{X\sim P}[-\log Q(X)].

The expectation is taken under PP, while the logarithmic probabilities come from QQ. Thus the order of the arguments matters. Base-two logarithms give units of bits; natural logarithms give nats. Terms with P(x)=0P(x)=0 contribute zero by convention. If P(x)>0P(x)>0 but Q(x)=0Q(x)=0, cross-entropy is infinite: the model declares an outcome impossible although the reference distribution allows it. (deeplearningbook.org)

The quantity −log⁡Q(x)-\log Q(x) is the self-information assigned to outcome xx by QQ. Cross-entropy therefore averages model-assigned surprise rather than the reference distribution’s own surprise. For discrete distributions it is nonnegative, but it need not vanish when the distributions agree. (deeplearningbook.org)

Relationship to entropy and divergence

Cross-entropy decomposes into Shannon entropy and Kullback–Leibler divergence:

H(P,Q)=H(P)+DKL(P∥Q),H(P,Q)=H(P)+D_{\mathrm{KL}}(P\|Q),

where

H(P)=−∑xP(x)log⁡P(x),DKL(P∥Q)=∑xP(x)log⁡P(x)Q(x).H(P)=-\sum_xP(x)\log P(x),\qquad D_{\mathrm{KL}}(P\|Q)=\sum_xP(x)\log\frac{P(x)}{Q(x)}.

For finite discrete distributions, Gibbs’ inequality implies H(P,Q)≥H(P)H(P,Q)\geq H(P), with equality precisely when P=QP=Q. Consequently, minimizing cross-entropy over QQ, with PP fixed, is equivalent to minimizing this direction of KL divergence. Cross-entropy is not a distance metric: it is generally asymmetric, and H(P,P)=H(P)H(P,P)=H(P), rather than zero. (cs229.stanford.edu)

In lossless data compression, ideal code lengths based on QQ are −log⁡2Q(x)-\log_2Q(x). Their average under the actual source PP is cross-entropy, while KL divergence represents the excess over entropy. This interpretation concerns ideal lengths or asymptotic coding rates; individual binary codewords must have integer lengths. (cs229.stanford.edu)

Statistical estimation and learning

In supervised learning, a model assigns conditional probabilities qθ(y∣x)q_\theta(y\mid x) to labels given inputs. For NN examples in training data, the empirical objective is

L(θ)=−1N∑i=1Nlog⁡qθ(yi∣xi).L(\theta)=-\frac1N\sum_{i=1}^{N} \log q_\theta(y_i\mid x_i).

When labels are conditionally independent across examples, their likelihood is the product of these probabilities. Taking its logarithm turns the product into a sum, so minimizing LL is equivalent to maximum likelihood estimation. The average loss is an empirical estimate of expected predictive logarithmic loss. (deeplearningbook.org)

Cross-entropy can serve as the training objective for an artificial neural network. The overall objective may also include a regularization term; in that case it is no longer solely the unmodified negative log-likelihood. Its precise expression depends on the output distribution assumed by the model. (deeplearningbook.org)

Binary and multiclass forms

For a binary target y∈{0,1}y\in\{0,1\} and predicted positive-class probability qq, the Bernoulli cross-entropy is

ℓ(y,q)=−ylog⁡q−(1−y)log⁡(1−q).\ell(y,q)=-y\log q-(1-y)\log(1-q).

The same formula accepts soft targets y∈[0,1]y\in[0,1], interpreted as target probabilities. In multilabel classification, separate binary losses can be calculated for labels that may coexist, rather than forcing all labels into one mutually exclusive distribution. (docs.pytorch.org)

For KK mutually exclusive classes, target probabilities pkp_k and predictions qkq_k give

ℓ(p,q)=−∑k=1Kpklog⁡qk.\ell(p,q)=-\sum_{k=1}^{K}p_k\log q_k.

With one-hot encoding of the correct class cc, this reduces to −log⁡qc-\log q_c. Consequently, only the probability assigned to the correct class appears explicitly, although normalization couples all class probabilities. Implementations may accept a class index instead of an explicit one-hot vector. (docs.pytorch.org)

For example, assigning probability 0.80.8 to the correct class produces approximately 0.2230.223 nats of loss; assigning 0.10.1 produces approximately 2.3032.303. These values follow directly from the formula and illustrate its strong penalty for confidently incorrect predictions.

Numerical computation

Multiclass models commonly transform raw scores zkz_k into probabilities using the softmax function:

qk=ezk∑jezj.q_k=\frac{e^{z_k}}{\sum_j e^{z_j}}.

For a normalized target distribution, differentiating cross-entropy with respect to a score gives

∂ℓ∂zk=qk−pk.\frac{\partial\ell}{\partial z_k}=q_k-p_k.

This compact expression supplies an output-layer gradient for backpropagation. It follows algebraically from the softmax and loss formulas. (docs.pytorch.org)

Computing logarithmic probabilities directly from scores avoids unnecessary numerical instability. Multiclass loss can be written using log-sum-exp, while binary implementations can combine the logistic function with the logarithmic loss. Libraries therefore distinguish losses receiving probabilities from those receiving raw scores. Class weights, averaging conventions, and ignored targets also affect the resulting objective. (docs.pytorch.org)

Language modeling and continuous distributions

A language model assigns conditional probabilities to successive tokens. Average negative log-probability on an independent test set estimates predictive cross-entropy per token. Exponentiating this average gives perplexity, using the same logarithm base. Comparisons require compatible token units and evaluation data. (nlp.stanford.edu)

For continuous variables with densities pp and qq, the analogous definition is

H(p,q)=−∫p(x)log⁡q(x) dx.H(p,q)=-\int p(x)\log q(x)\,dx.

Unlike discrete cross-entropy, this quantity can be negative because a density can exceed one. It relates to differential entropy through the corresponding KL decomposition when the quantities are well-defined; its value depends on the chosen coordinates and reference measure. (deeplearningbook.org)