Marginal likelihood is a quantity in Bayesian inference that measures the probability, or probability density, of observed data under a statistical model after averaging over uncertain parameters. Also called model evidence, it combines a model’s likelihood function with its prior distribution. It normalizes the parameter posterior and provides the basis for comparing models through Bayes factors. Unlike a maximized likelihood, it evaluates the model’s predictions across its parameter space rather than only at its best-fitting parameter values. (pmc.ncbi.nlm.nih.gov)
Definition and interpretation
Let denote observed data, a model, and its parameters. For continuous parameters, the marginal likelihood is
The integration removes, or marginalizes, from the joint distribution of data and parameters. For discrete parameters, the integral becomes a sum; mixed parameter spaces require both operations. Equivalently, marginal likelihood is the prior expectation of the likelihood:
Viewed as a function of possible datasets, it is the model’s prior predictive distribution. For continuous observations, it is a density, not the probability of observing an exact data vector, and can exceed one. (arxiv.org)
By Bayes’ theorem, the posterior distribution is
provided the denominator is finite and positive. The evidence is therefore the normalizing constant that converts the likelihood–prior product into a probability distribution. (gaussianprocess.org)
Model comparison and complexity
For models and , their Bayes factor is
Posterior model odds equal this ratio multiplied by prior model odds. Thus, evidence alone is not a model’s posterior probability. With multiple candidate models,
These probabilities can also supply weights for Bayesian model averaging, which incorporates uncertainty about model structure into estimation and prediction. Competing models need not be nested. (stat.cmu.edu)
Marginal likelihood expresses a trade-off between fit and predictive breadth. A flexible model may fit many possible datasets, but its prior predictive probability must be distributed among them. High likelihood confined to a small region of prior probability can therefore contribute less evidence than moderately high likelihood across a substantial region. This is often described as a Bayesian Occam factor. It is not simply a penalty proportional to parameter count, nor a guarantee that the simplest model wins. Unlike maximum likelihood estimation, it averages rather than maximizes over parameters. (gaussianprocess.org)
A conjugate example
Consider Bernoulli observations with successes, conditionally independent given a success probability . Assign a beta distribution with positive shape parameters as a conjugate prior. For a particular ordered sequence,
Direct integration gives
where is the beta function. If the recorded observation is instead the success count , the binomial sampling likelihood includes the number of sequences yielding that count:
The distinction illustrates that evidence depends on precisely what constitutes the observed data. Common factors may cancel in a Bayes factor, but remain part of each marginal likelihood. (pmc.ncbi.nlm.nih.gov)
Computation and approximation
Closed-form evidence is available for some conjugate models, but high-dimensional integration is frequently difficult. Numerical quadrature is practical in sufficiently low dimensions. The Laplace approximation replaces a locally concentrated posterior with a Gaussian approximation, using curvature near a mode. Its accuracy can deteriorate for strongly skewed or multimodal distributions. (arxiv.org)
Simulation methods include importance sampling, bridge sampling, and nested sampling. Bridge sampling combines posterior draws with draws from an auxiliary distribution to estimate the normalizing constant. Ordinary Markov chain Monte Carlo can generate posterior samples without evaluating evidence, so successful posterior sampling does not automatically supply an evidence estimate. Simple harmonic-mean estimators can be highly unstable. (pmc.ncbi.nlm.nih.gov)
Nested sampling reformulates evidence calculation through the prior probability mass enclosed by likelihood thresholds. It estimates evidence while also producing information about the posterior; sampling accurately from likelihood-restricted priors is a central computational challenge. (arxiv.org)
In variational inference, an approximating distribution yields an evidence lower bound:
Because KL divergence is nonnegative, the ELBO cannot exceed log evidence. Comparing bounds is not necessarily equivalent to comparing exact evidences, because approximation gaps can differ between models. (cs.columbia.edu)
Prior dependence and machine learning
Evidence depends on the normalized parameter prior. Broadening a prior can reduce evidence by allocating more probability to poorly fitting parameter values. An improper prior generally leaves evidence undefined up to an arbitrary multiplicative constant, even when the posterior is proper. Consequently, ordinary evidence-based model comparison requires careful specification of proper priors, apart from special constructions where ambiguities are resolved. (stat.cmu.edu)
In machine learning, evidence can be optimized over a hyperparameter while integrating out lower-level parameters or latent quantities. This is commonly called empirical Bayes or type-II maximum likelihood. In Gaussian process regression with Gaussian observation noise, the marginal likelihood is analytically available and is used to estimate covariance and noise hyperparameters. Optimizing those hyperparameters is distinct from integrating over their uncertainty in a fully Bayesian analysis. (gaussianprocess.org)