Cross-entropy is a quantity in information theory that measures the average negative logarithm of probabilities assigned by a model distribution to outcomes generated by a reference distribution. It connects probabilistic prediction with information coding: predictions assigning low probability to frequently occurring outcomes incur a larger cost. In machine learning, cross-entropy is widely used as a loss function for fitting probabilistic models and evaluating their predictions. (deeplearningbook.org)
Mathematical definition
For two discrete probability distributions, and , over the same outcome space , cross-entropy is
The expectation is taken under , while the logarithmic probabilities come from . Thus the order of the arguments matters. Base-two logarithms give units of bits; natural logarithms give nats. Terms with contribute zero by convention. If but , cross-entropy is infinite: the model declares an outcome impossible although the reference distribution allows it. (deeplearningbook.org)
The quantity is the self-information assigned to outcome by . Cross-entropy therefore averages model-assigned surprise rather than the reference distribution’s own surprise. For discrete distributions it is nonnegative, but it need not vanish when the distributions agree. (deeplearningbook.org)
Relationship to entropy and divergence
Cross-entropy decomposes into Shannon entropy and Kullback–Leibler divergence:
where
For finite discrete distributions, Gibbs’ inequality implies , with equality precisely when . Consequently, minimizing cross-entropy over , with fixed, is equivalent to minimizing this direction of KL divergence. Cross-entropy is not a distance metric: it is generally asymmetric, and , rather than zero. (cs229.stanford.edu)
In lossless data compression, ideal code lengths based on are . Their average under the actual source is cross-entropy, while KL divergence represents the excess over entropy. This interpretation concerns ideal lengths or asymptotic coding rates; individual binary codewords must have integer lengths. (cs229.stanford.edu)
Statistical estimation and learning
In supervised learning, a model assigns conditional probabilities to labels given inputs. For examples in training data, the empirical objective is
When labels are conditionally independent across examples, their likelihood is the product of these probabilities. Taking its logarithm turns the product into a sum, so minimizing is equivalent to maximum likelihood estimation. The average loss is an empirical estimate of expected predictive logarithmic loss. (deeplearningbook.org)
Cross-entropy can serve as the training objective for an artificial neural network. The overall objective may also include a regularization term; in that case it is no longer solely the unmodified negative log-likelihood. Its precise expression depends on the output distribution assumed by the model. (deeplearningbook.org)
Binary and multiclass forms
For a binary target and predicted positive-class probability , the Bernoulli cross-entropy is
The same formula accepts soft targets , interpreted as target probabilities. In multilabel classification, separate binary losses can be calculated for labels that may coexist, rather than forcing all labels into one mutually exclusive distribution. (docs.pytorch.org)
For mutually exclusive classes, target probabilities and predictions give
With one-hot encoding of the correct class , this reduces to . Consequently, only the probability assigned to the correct class appears explicitly, although normalization couples all class probabilities. Implementations may accept a class index instead of an explicit one-hot vector. (docs.pytorch.org)
For example, assigning probability to the correct class produces approximately nats of loss; assigning produces approximately . These values follow directly from the formula and illustrate its strong penalty for confidently incorrect predictions.
Numerical computation
Multiclass models commonly transform raw scores into probabilities using the softmax function:
For a normalized target distribution, differentiating cross-entropy with respect to a score gives
This compact expression supplies an output-layer gradient for backpropagation. It follows algebraically from the softmax and loss formulas. (docs.pytorch.org)
Computing logarithmic probabilities directly from scores avoids unnecessary numerical instability. Multiclass loss can be written using log-sum-exp, while binary implementations can combine the logistic function with the logarithmic loss. Libraries therefore distinguish losses receiving probabilities from those receiving raw scores. Class weights, averaging conventions, and ignored targets also affect the resulting objective. (docs.pytorch.org)
Language modeling and continuous distributions
A language model assigns conditional probabilities to successive tokens. Average negative log-probability on an independent test set estimates predictive cross-entropy per token. Exponentiating this average gives perplexity, using the same logarithm base. Comparisons require compatible token units and evaluation data. (nlp.stanford.edu)
For continuous variables with densities and , the analogous definition is
Unlike discrete cross-entropy, this quantity can be negative because a density can exceed one. It relates to differential entropy through the corresponding KL decomposition when the quantities are well-defined; its value depends on the chosen coordinates and reference measure. (deeplearningbook.org)