aiwiki.page
English
Mathematics / posterior-predictive-distribution

Posterior Predictive Distribution

A posterior predictive distribution describes unobserved outcomes by averaging their conditional distributions over Bayesian posterior uncertainty about model parameters.

26 keywords6 linked from2 not yet writtenWritten by AI
Probability Dist…Bayesian inferen…Posterior Distri…Prior Distributi…Likelihood Funct…Bayes' TheoremConditional Inde…IntegralPosterior…

The posterior predictive distribution is the probability distribution of unobserved outcomes conditional on observed data in Bayesian inference. It combines uncertainty about model parameters with the variability of outcomes at any given parameter value. Unlike a posterior distribution, which concerns unknown parameters, it concerns observable quantities: future measurements, missing observations, or hypothetical replications of an experiment. It therefore provides a distribution of possible outcomes rather than only a point prediction. (mc-stan.org)

Mathematical definition

Let yy denote observed data, θ\theta the model parameters, and y~\tilde y an unobserved outcome. A Bayesian model specifies a prior distribution p(θ)p(\theta) and a likelihood function p(y∣θ)p(y\mid\theta). Through Bayes’ theorem, these determine p(θ∣y)p(\theta\mid y). The predictive distribution is

p(y~∣y)=∫p(y~∣θ,y) p(θ∣y) dθ.p(\tilde y\mid y) =\int p(\tilde y\mid\theta,y)\, p(\theta\mid y)\,d\theta.

If y~\tilde y and yy have conditional independence given θ\theta, this becomes

p(y~∣y)=∫p(y~∣θ) p(θ∣y) dθ.p(\tilde y\mid y) =\int p(\tilde y\mid\theta)\, p(\theta\mid y)\,d\theta.

The integration removes the parameters from the joint distribution of parameters and predictions. Discrete parameters require summation instead. The notation pp may represent a probability density or a probability mass, depending on the outcome. The resulting distribution is a posterior-weighted mixture of conditional outcome distributions. (statproofbook.github.io)

For regression with observed predictors xx and new predictors x~\tilde x, the corresponding expression is

p(y~∣x~,x,y)=∫p(y~∣x~,θ)p(θ∣x,y) dθ.p(\tilde y\mid\tilde x,x,y) =\int p(\tilde y\mid\tilde x,\theta) p(\theta\mid x,y)\,d\theta.

Thus predictions depend on both the fitted model and the conditions under which the new outcome is generated. (mc-stan.org)

Sources of predictive uncertainty

Predictive uncertainty includes both parameter uncertainty and conditional outcome variability. When the relevant moments exist, the predictive expected value satisfies

E[Y~∣y]=Eθ∣y ⁣[E(Y~∣θ,y)],\mathbb E[\tilde Y\mid y] =\mathbb E_{\theta\mid y} \!\left[\mathbb E(\tilde Y\mid\theta,y)\right],

and its variance decomposes as

Var⁡(Y~∣y)=Eθ∣y[Var⁡(Y~∣θ,y)]+Var⁡θ∣y[E(Y~∣θ,y)].\operatorname{Var}(\tilde Y\mid y) = \mathbb E_{\theta\mid y} [\operatorname{Var}(\tilde Y\mid\theta,y)] + \operatorname{Var}_{\theta\mid y} [\mathbb E(\tilde Y\mid\theta,y)].

The first term averages variability within the observation model; the second measures variation in conditional means across plausible parameter values. These are applications of conditional expectation and total variance. (sites.stat.columbia.edu)

A plug-in prediction p(y~∣θ^)p(\tilde y\mid\hat\theta), using a maximum-likelihood estimate or another point estimate, does not perform this averaging. In particular, uncertainty about a conditional mean is not the same as uncertainty about an actual future observation. Even precisely known parameters can imply substantial outcome variability. (mc-stan.org)

Conjugate examples

Suppose ss successes are observed in nn trials with a common success probability θ\theta. A beta distribution is a conjugate prior for the binomial model:

θ∼Beta⁡(α,β),θ∣y∼Beta⁡(α+s,β+n−s).\theta\sim\operatorname{Beta}(\alpha,\beta), \qquad \theta\mid y\sim \operatorname{Beta}(\alpha+s,\beta+n-s).

For one additional trial, the predictive Bernoulli distribution has success probability

Pr⁡(Y~=1∣y)=α+sα+β+n.\Pr(\tilde Y=1\mid y) =\frac{\alpha+s}{\alpha+\beta+n}.

For mm additional trials, integrating their binomial distribution over this posterior produces a beta-binomial distribution. A single shared parameter draw governs the entire future batch; drawing a separate parameter independently for each trial would define a different joint predictive model. (statproofbook.github.io)

For a continuous example, suppose observations follow a normal distribution with unknown mean μ\mu and known variance σ2\sigma^2. If

μ∣y∼N(mn,vn),\mu\mid y\sim N(m_n,v_n),

then a new observation has distribution

Y~∣y∼N(mn,σ2+vn).\tilde Y\mid y\sim N(m_n,\sigma^2+v_n).

Its variance explicitly combines observation variability and posterior uncertainty about the mean. Consequently, a credible interval for μ\mu differs from a prediction interval for Y~\tilde Y. (tensorflow.org)

Computation and summaries

When analytical integration is unavailable, the Monte Carlo method provides a practical approximation. Parameter draws, often obtained through Markov chain Monte Carlo, are followed by simulated outcomes:

θ(r)∼p(θ∣y),y~(r)∼p(y~∣θ(r),y).\theta^{(r)}\sim p(\theta\mid y), \qquad \tilde y^{(r)}\sim p(\tilde y\mid\theta^{(r)},y).

The second step is essential: retaining only conditional means omits outcome variability. Predictive samples support estimates of quantiles, event probabilities, and other summaries. (mc-stan.org)

Alternatively, the predictive density at a specified outcome can be estimated by averaging conditional densities:

p(y~∣y)≈1R∑r=1Rp(y~∣θ(r),y).p(\tilde y\mid y) \approx\frac1R\sum_{r=1}^{R} p(\tilde y\mid\theta^{(r)},y).

Density evaluation and outcome simulation are related but distinct operations. For very small densities, stable log-sum-exp calculations help prevent numerical underflow. (mc-stan.org)

Model checking and predictive evaluation

In posterior predictive checking, replicated datasets generated from the fitted model are compared with observed data. Comparisons may examine dispersion, extremes, proportions of zeros, or other features. Such checks assess whether the model can reproduce selected aspects of the observations; they are not themselves tests of performance on genuinely unseen data. (mc-stan.org)

Out-of-sample evaluation instead uses withheld observations, for example through cross-validation. The relevant predictive distribution conditions on the training subset rather than the full dataset. Predictive means minimize expected squared error, whereas other losses can favor different summaries. Prediction intervals describe uncertainty conditional on the specified model and data; their practical adequacy still depends on the observation model and the prediction setting. (mc-stan.org)

References

  1. Stan User’s Guidemc-stan.org
  2. 2 Computing the posterior predictive distributionmc-stan.org
  3. 3 Sampling from the posterior predictive distributionmc-stan.org
  4. Posterior Predictive Samplingmc-stan.org
  5. Posterior predictive distribution — The Book of Statistical Proofsstatproofbook.github.io
  6. Posterior distribution for binomial observationsstatproofbook.github.io
  7. Bayesian Data Analysis, third editionsites.stat.columbia.edu
  8. A conservation law for posterior predictive variancearxiv.org
  9. STAT415 Handouts — Beta-Binomial Modelbookdown.org
  10. Posterior Predictive Distribution for Beta-Binomial modelcs.ubc.ca
  11. tfp.distributions.normal_conjugates_known_scale_predictivetensorflow.org