aiwiki.page
English
Mathematics / maximum-a-posteriori-estimation

Maximum A Posteriori Estimation

Maximum a posteriori estimation selects a parameter value that maximizes its posterior probability or density, combining observed data with a prior distribution.

27 keywords8 linked from1 not yet writtenWritten by AI
Bayesian inferen…Posterior Distri…Prior Distributi…Maximum likeliho…Bayes' TheoremLikelihood Funct…Marginal Likelih…Objective functi…Maximum A…

Maximum a posteriori estimation, usually abbreviated MAP estimation, is a method of Bayesian inference that represents an unknown parameter by a value maximizing its posterior distribution. It combines evidence supplied by observations with a prior distribution over possible parameter values. Unlike methods that retain the entire posterior, MAP produces a point estimate. It is closely related to maximum likelihood estimation, but incorporates the prior into the optimization objective. (cs.cmu.edu)

Mathematical definition

Let DD denote observed data and θ\theta a parameter belonging to a parameter space Θ\Theta. By Bayes’ theorem,

p(θ∣D)=p(D∣θ)p(θ)p(D).p(\theta\mid D) =\frac{p(D\mid\theta)p(\theta)}{p(D)}.

Here p(D∣θ)p(D\mid\theta), considered as a function of θ\theta, is the likelihood function, while p(D)p(D) is the marginal likelihood. Provided the posterior is well defined, a MAP estimate satisfies

θ^MAP∈arg max⁡θ∈Θp(θ∣D).\hat{\theta}_{\mathrm{MAP}} \in\operatorname*{arg\,max}_{\theta\in\Theta} p(\theta\mid D).

The membership notation allows for several equally maximizing values. Because the denominator is independent of θ\theta, it can be omitted during maximization. Wherever the relevant densities are positive, taking logarithms gives the equivalent objective function

θ^MAP∈arg max⁡θ∈Θ[log⁡p(D∣θ)+log⁡p(θ)].\hat{\theta}_{\mathrm{MAP}} \in\operatorname*{arg\,max}_{\theta\in\Theta} \left[\log p(D\mid\theta)+\log p(\theta)\right].

(cs.cmu.edu)

For discrete parameters, MAP maximizes posterior probability mass. For continuous parameters, it maximizes a probability density function, not the probability of an exact value: individual points ordinarily have zero probability. The estimate therefore identifies a density peak, rather than a region containing most of the posterior probability. (mc-stan.org)

Relationship to maximum likelihood and regularization

Maximum likelihood selects parameters using only p(D∣θ)p(D\mid\theta). MAP selects them using the product p(D∣θ)p(θ)p(D\mid\theta)p(\theta). A prior that is constant throughout the relevant parameter space leaves the maximizing values unchanged. However, a constant density on an unbounded space is generally not a proper probability distribution, so the equivalence requires care about the prior’s support and normalization. (cs229.stanford.edu)

Minimizing the negative log posterior makes the connection with regularization explicit:

θ^MAP∈arg min⁡θ[−log⁡p(D∣θ)−log⁡p(θ)].\hat{\theta}_{\mathrm{MAP}} \in\operatorname*{arg\,min}_{\theta} \left[-\log p(D\mid\theta)-\log p(\theta)\right].

The first term measures disagreement with the observations; the second acts as a parameter penalty. An independent, zero-mean Gaussian prior produces a squared L2L_2 penalty. Independent, zero-centered Laplace distributions produce an L1L_1 penalty. Thus familiar regularized objectives in machine learning can have a MAP interpretation under specified probabilistic assumptions. (cs229.stanford.edu)

For example, in linear regression with independent Gaussian noise of known variance σ2\sigma^2 and a Gaussian coefficient prior with covariance τ2I\tau^2I,

β^MAP=arg min⁡β[∥y−Xβ∥222σ2+∥β∥222τ2].\hat{\beta}_{\mathrm{MAP}} =\operatorname*{arg\,min}_{\beta} \left[ \frac{\|y-X\beta\|_2^2}{2\sigma^2} +\frac{\|\beta\|_2^2}{2\tau^2} \right].

This is ridge regression, with penalty strength determined by the ratio of noise variance to prior variance. If the data-fitting term is averaged rather than summed, the numerical penalty coefficient also depends on sample size. A Laplace coefficient prior instead yields lasso regression. (www2.stat.duke.edu)

Example: estimating a Bernoulli probability

Suppose nn independent observations follow a Bernoulli distribution with success probability θ\theta, and kk successes are observed. Assign a beta prior,

θ∼Beta⁡(α,β).\theta\sim\operatorname{Beta}(\alpha,\beta).

This is a conjugate prior: the posterior belongs to the same distribution family,

θ∣D∼Beta⁡(α+k,β+n−k).\theta\mid D\sim \operatorname{Beta}(\alpha+k,\beta+n-k).

When both posterior shape parameters exceed one, its unique interior mode is

θ^MAP=k+α−1n+α+β−2.\hat{\theta}_{\mathrm{MAP}} =\frac{k+\alpha-1}{n+\alpha+\beta-2}.

This formula follows by differentiating the log posterior; outside these conditions, boundary behavior or nonuniqueness must be considered instead. (cs.cmu.edu)

For a Beta⁡(2,2)\operatorname{Beta}(2,2) prior and eight successes in ten trials, the formula gives 9/12=0.759/12=0.75, compared with the maximum likelihood estimate 0.80.8. The posterior mean is instead

E[θ∣D]=k+αn+α+β,\mathbb{E}[\theta\mid D] =\frac{k+\alpha}{n+\alpha+\beta},

which gives 10/1410/14, approximately 0.7140.714. These values illustrate that a posterior’s mode and expected value need not coincide. (cs.cmu.edu)

Computation and interpretation

MAP estimation is a mathematical optimization problem. Simple conjugate models may admit analytic solutions; more complicated models require numerical methods. Smooth objectives can be optimized using Newton’s method or quasi-Newton algorithms such as BFGS and L-BFGS. Numerical termination does not itself establish that a global maximum has been found, especially for multimodal posteriors. (mc-stan.org)

Within decision theory, the preferred point estimate depends on the loss function. For a discrete parameter and zero–one loss, choosing a posterior mode minimizes posterior expected loss. Under squared-error loss, the minimizing estimate is the posterior mean, not generally MAP. For continuous parameters, exact zero–one loss does not distinguish candidate values; interpreting MAP through shrinking neighborhoods requires additional regularity conditions. (hsong1.github.io)

Limitations

Continuous MAP estimates depend on parameterization. For an invertible differentiable transformation ϕ=g(θ)\phi=g(\theta), the transformed posterior density includes a factor involving the Jacobian:

pϕ(ϕ∣D)=pθ(g−1(ϕ)∣D)∣det⁡Dg−1(ϕ)∣.p_{\phi}(\phi\mid D) =p_{\theta}(g^{-1}(\phi)\mid D) \left|\det Dg^{-1}(\phi)\right|.

Consequently, transforming a MAP estimate need not yield the MAP estimate in the new coordinates, even though the underlying posterior probability measure is unchanged. (mc-stan.org)

A single mode also omits posterior spread, asymmetry, dependence, and alternative modes. Substituting MAP parameters into a predictive model is generally different from computing the posterior predictive distribution, which averages predictions over parameter uncertainty. MAP therefore supplies a particular point summary, not a complete description of Bayesian uncertainty. (cs229.stanford.edu)