A conjugate prior is a prior distribution that, under a specified statistical model, produces a posterior distribution belonging to the same family of probability distributions. In Bayesian inference, this property makes updating especially convenient: observations change the distribution’s parameters without changing its functional form. Conjugacy is a relationship between a prior family and a likelihood function, not an intrinsic property of a distribution considered alone. (math.mit.edu)
Definition and mathematical structure
Let denote observed data, an unknown parameter, and a prior density indexed by hyperparameters . By Bayes’ theorem,
The denominator is the marginal likelihood. The prior family is conjugate if every admissible update can be written as , where depends on the observations and original hyperparameters. Thus, posterior calculation often reduces to multiplying kernels and recognizing a familiar probability density function. The prior and posterior need not have the same parameter values, and neither must belong to the distribution family used for the observations. (stat210a.berkeley.edu)
Beta–binomial conjugacy
Suppose , the number of successes in trials, follows a binomial distribution with unknown success probability . Its likelihood is proportional to
A beta distribution with positive parameters has density proportional to . Multiplying these expressions gives
The same update applies to a sequence of conditionally independent Bernoulli observations. The prior hyperparameters are updated by adding successes and failures, respectively. (ocw.mit.edu)
The posterior mean is
for . This expresses the estimate as a weighted average of the prior mean and observed success proportion. In this mean-based interpretation, acts as an effective prior sample size. Hyperparameters need not be integers or represent actual previous observations. (stat210a.berkeley.edu)
For example, a prior and seven successes in ten trials give a posterior, with mean , approximately . This is a direct application of the update formula. (ocw.mit.edu)
Other standard conjugate pairs
For conditionally independent counts from a Poisson distribution with rate , a gamma prior is conjugate. Using shape and rate , its kernel is , and
The rate convention matters: formulations using a scale parameter require different-looking updates. (math.mit.edu)
For observations from a normal distribution with unknown mean and known variance , a normal prior yields
Posterior precision—the reciprocal of variance—is the sum of prior and data precisions. When both mean and variance are unknown, standard joint conjugate families include the normal–inverse-gamma distribution, rather than an arbitrary pair of independent priors. (cs.ubc.ca)
Connection with exponential families
A broad construction is available for exponential-family models. Write a sampling density as
where is the natural parameter, a statistic, and the log-normalizing function. For independent observations, is a sufficient statistic. A conjugate prior density over , relative to a specified reference measure, has kernel
Multiplication by the likelihood gives the additive updates
These formulas apply where the prior and posterior normalizing integrals are finite. They explain why many familiar conjugate updates accumulate counts, sums, or other low-dimensional statistics. However, the construction does not guarantee an elementary expression for the normalizing constant. (stat.berkeley.edu)
Prediction and sequential updating
Conjugacy concerns uncertainty about parameters, whereas the posterior predictive distribution describes future observations:
Integrating a binomial sampling distribution against a beta posterior produces a beta–binomial distribution. The predictive family therefore need not match either the posterior family or the original sampling family. (lancaster.ac.uk)
With observations independent conditional on the parameter, updating sequentially is equivalent to updating with the entire dataset at once. The posterior after one batch becomes the prior for the next; additive sufficient statistics make such updates compact. (ocw.mit.edu)
Scope and limitations
Conjugate families are not unique: mixtures of conjugate priors can also remain closed under updating, with both component parameters and mixture weights changing. Their mathematical convenience does not establish that their shapes adequately represent substantive prior information. Conjugacy also does not guarantee computational tractability when normalizing constants are difficult to evaluate. (stat.berkeley.edu)
In multiparameter models, convenient full conditional distributions may permit Gibbs sampling, a Markov chain Monte Carlo method, even when directly sampling the joint posterior is difficult. Such conditional updating is distinct from obtaining a complete joint posterior through a single conjugate-family update. (sites.stat.columbia.edu)