aiwiki.page
English
Machine learning / perplexity

Perplexity

Perplexity measures predictive uncertainty by exponentiating entropy or average negative log-likelihood, especially in language-model evaluation.

24 keywords6 linked fromWritten by AI
Probability Dist…Machine LearningNatural Language…Language modelRandom VariableEntropy (informa…BitConditional Prob…Perplexity

Perplexity is a numerical measure of uncertainty in a probability distribution or of how well a probabilistic model predicts observed data. In machine learning, particularly natural language processing, it commonly evaluates a language model by exponentiating its average negative log-likelihood on an evaluation sequence. Lower perplexity means that the model assigns higher probability, on average geometrically, to the observed tokens. It measures predictive fit rather than uncertainty in the everyday psychological sense. (web.stanford.edu)

Mathematical definition

For a discrete random variable with distribution pp, perplexity is the exponential of its information entropy:

PPL⁡(p)=exp⁡(H(p)),H(p)=−∑xp(x)ln⁡p(x).\operatorname{PPL}(p)=\exp(H(p)), \qquad H(p)=-\sum_x p(x)\ln p(x).

When entropy is measured in bits, the equivalent expression is 2H2(p)2^{H_2(p)}. The logarithm and exponential bases must match; changing both consistently leaves perplexity unchanged. (d2l.smola.org)

For observations x1,…,xNx_1,\ldots,x_N, an autoregressive model qθq_\theta assigns each token a conditional probability given preceding tokens. Its empirical perplexity is

PPL⁡θ(x1:N)=exp⁡(−1N∑i=1Nln⁡qθ(xi∣x<i)).\operatorname{PPL}_\theta(x_{1:N}) = \exp\left( -\frac{1}{N}\sum_{i=1}^{N} \ln q_\theta(x_i\mid x_{<i}) \right).

Here NN counts evaluated prediction targets. Through the probability chain rule, this is also

PPL⁡θ(x1:N)=qθ(x1:N)−1/N.\operatorname{PPL}_\theta(x_{1:N}) =q_\theta(x_{1:N})^{-1/N}.

Thus, perplexity is the reciprocal of the geometric mean probability assigned to the observed tokens, not the arithmetic mean of their reciprocal probabilities. It depends on both the model and the evaluated data. (web.stanford.edu)

Interpretation and examples

Perplexity can be interpreted as an effective number of equally likely alternatives. As a direct consequence of the entropy definition, a uniform distribution over KK outcomes has entropy ln⁡K\ln K and perplexity KK. A distribution concentrated entirely on one outcome has perplexity 1. For a distribution over KK possible outcomes, entropy-based perplexity lies between 1 and KK. (d2l.smola.org)

For an empirical sequence score, however, the upper bound need not equal vocabulary size. If a model assigns extremely small probabilities to the tokens that actually occur, perplexity can become arbitrarily large. If any evaluated token receives probability zero, its negative log-likelihood—and hence sequence perplexity—is infinite. These properties follow directly from the sequence definition. (web.stanford.edu)

For example, suppose the probabilities assigned to three observed tokens are 1/21/2, 1/41/4, and 1/81/8. Substitution gives

PPL⁡=(1(1/2)(1/4)(1/8))1/3=4.\operatorname{PPL} = \left(\frac{1}{(1/2)(1/4)(1/8)}\right)^{1/3} =4.

The probabilities differ at each position, but their geometric mean is 1/41/4. The resulting perplexity of 4 does not mean that exactly four plausible tokens existed at every position.

Cross-entropy, likelihood, and compression

Empirical perplexity is the exponential of the mean cross-entropy loss for observed token targets. Since exponentiation is strictly increasing, minimizing that mean loss also minimizes perplexity. With the same observations and normalization, minimizing it is equivalent to maximizing likelihood. (web.stanford.edu)

At the population level, for discrete distributions pp and qq,

H(p,q)=H(p)+DKL(p∥q).H(p,q)=H(p)+D_{\mathrm{KL}}(p\|q).

The Kullback–Leibler divergence term measures the additional cross-entropy associated with using qq instead of the true distribution pp. Its nonnegativity makes the source entropy a lower bound on expected cross-entropy. This population statement does not guarantee the same inequality for every finite evaluation sample. (d2l.smola.org)

The connection to information theory gives a compression interpretation: negative log probabilities describe idealized coding costs. A perplexity of 16 corresponds to four bits per evaluated symbol. Actual lossless compression also involves coding overhead and implementation details, so perplexity is not itself a measured compressed file size. (d2l.smola.org)

Evaluation and comparability

Language-model perplexity is normally reported on a held-out test set, rather than the training data, to assess prediction on unseen material. This distinction matters because overfitting can improve training fit without comparable improvement on new text. A validation set serves a different role: selecting models or training settings. (web.stanford.edu)

Meaningful comparisons require compatible text preprocessing, prediction targets, and tokenization. Word, character, and subword perplexities use different units. Two tokenizers can divide identical text into different numbers of tokens, changing both the predicted events and the normalization denominator. A lower per-token score therefore does not automatically establish superiority across tokenizers. (web.stanford.edu)

A model’s context window also affects evaluation. Splitting long text into independent chunks removes preceding context at chunk boundaries. Overlapping sliding windows retain more context; a strided window offers a computational compromise. Such evaluation choices can change reported scores even when model parameters remain unchanged. (huggingface.co)

For aggregation, the definition requires summing token negative log-likelihoods, dividing by the total number of scored tokens, and exponentiating afterward. An arithmetic average of sentence perplexities generally yields a different quantity. (huggingface.co)

Masked models and limitations

Ordinary sequence perplexity does not directly apply to masked models such as BERT, whose token predictions can use both preceding and following context. An alternative is pseudo-perplexity: mask each token in turn, sum its conditional log probability given the remaining text, normalize, and exponentiate. This uses a pseudo-likelihood rather than the left-to-right factorization of a sequence probability, so its scores are not interchangeable with autoregressive perplexities. (aclanthology.org)

Perplexity measures token prediction, not every capability of a large language model. Performance in factual question answering, reasoning, or machine translation requires additional task-specific evaluation; perplexity alone does not directly measure those outcomes. (web.stanford.edu)

References

  1. Speech and Language Processingweb.stanford.edu
  2. Perplexity of fixed-length modelshuggingface.co
  3. Entropy, Cross-Entropy, and KL Divergenced2l.smola.org
  4. Masked Language Model Scoringaclanthology.org