The posterior predictive distribution is the probability distribution of unobserved outcomes conditional on observed data in Bayesian inference. It combines uncertainty about model parameters with the variability of outcomes at any given parameter value. Unlike a posterior distribution, which concerns unknown parameters, it concerns observable quantities: future measurements, missing observations, or hypothetical replications of an experiment. It therefore provides a distribution of possible outcomes rather than only a point prediction. (mc-stan.org)
Mathematical definition
Let denote observed data, the model parameters, and an unobserved outcome. A Bayesian model specifies a prior distribution and a likelihood function . Through Bayes’ theorem, these determine . The predictive distribution is
If and have conditional independence given , this becomes
The integration removes the parameters from the joint distribution of parameters and predictions. Discrete parameters require summation instead. The notation may represent a probability density or a probability mass, depending on the outcome. The resulting distribution is a posterior-weighted mixture of conditional outcome distributions. (statproofbook.github.io)
For regression with observed predictors and new predictors , the corresponding expression is
Thus predictions depend on both the fitted model and the conditions under which the new outcome is generated. (mc-stan.org)
Sources of predictive uncertainty
Predictive uncertainty includes both parameter uncertainty and conditional outcome variability. When the relevant moments exist, the predictive expected value satisfies
and its variance decomposes as
The first term averages variability within the observation model; the second measures variation in conditional means across plausible parameter values. These are applications of conditional expectation and total variance. (sites.stat.columbia.edu)
A plug-in prediction , using a maximum-likelihood estimate or another point estimate, does not perform this averaging. In particular, uncertainty about a conditional mean is not the same as uncertainty about an actual future observation. Even precisely known parameters can imply substantial outcome variability. (mc-stan.org)
Conjugate examples
Suppose successes are observed in trials with a common success probability . A beta distribution is a conjugate prior for the binomial model:
For one additional trial, the predictive Bernoulli distribution has success probability
For additional trials, integrating their binomial distribution over this posterior produces a beta-binomial distribution. A single shared parameter draw governs the entire future batch; drawing a separate parameter independently for each trial would define a different joint predictive model. (statproofbook.github.io)
For a continuous example, suppose observations follow a normal distribution with unknown mean and known variance . If
then a new observation has distribution
Its variance explicitly combines observation variability and posterior uncertainty about the mean. Consequently, a credible interval for differs from a prediction interval for . (tensorflow.org)
Computation and summaries
When analytical integration is unavailable, the Monte Carlo method provides a practical approximation. Parameter draws, often obtained through Markov chain Monte Carlo, are followed by simulated outcomes:
The second step is essential: retaining only conditional means omits outcome variability. Predictive samples support estimates of quantiles, event probabilities, and other summaries. (mc-stan.org)
Alternatively, the predictive density at a specified outcome can be estimated by averaging conditional densities:
Density evaluation and outcome simulation are related but distinct operations. For very small densities, stable log-sum-exp calculations help prevent numerical underflow. (mc-stan.org)
Model checking and predictive evaluation
In posterior predictive checking, replicated datasets generated from the fitted model are compared with observed data. Comparisons may examine dispersion, extremes, proportions of zeros, or other features. Such checks assess whether the model can reproduce selected aspects of the observations; they are not themselves tests of performance on genuinely unseen data. (mc-stan.org)
Out-of-sample evaluation instead uses withheld observations, for example through cross-validation. The relevant predictive distribution conditions on the training subset rather than the full dataset. Predictive means minimize expected squared error, whereas other losses can favor different summaries. Prediction intervals describe uncertainty conditional on the specified model and data; their practical adequacy still depends on the observation model and the prediction setting. (mc-stan.org)
References
- Stan User’s Guidemc-stan.org
- 2 Computing the posterior predictive distributionmc-stan.org
- 3 Sampling from the posterior predictive distributionmc-stan.org
- Posterior Predictive Samplingmc-stan.org
- Posterior predictive distribution — The Book of Statistical Proofsstatproofbook.github.io
- Posterior distribution for binomial observationsstatproofbook.github.io
- Bayesian Data Analysis, third editionsites.stat.columbia.edu
- A conservation law for posterior predictive variancearxiv.org
- STAT415 Handouts — Beta-Binomial Modelbookdown.org
- Posterior Predictive Distribution for Beta-Binomial modelcs.ubc.ca
- tfp.distributions.normal_conjugates_known_scale_predictivetensorflow.org