Variational inference (VI) is a family of methods that approximate difficult-to-compute probability distributions using mathematical optimization. In Bayesian inference, it replaces an intractable posterior distribution with a tractable distribution selected from a specified family. Rather than recovering the posterior through sampling alone, VI adjusts the approximation to optimize a measure of agreement with the target. It is used in statistics and machine learning, particularly when exact inference is computationally impractical. (cs.columbia.edu)
Mathematical formulation
Let denote observed data and denote unknown parameters or latent variables. By Bayes’ theorem,
The denominator is the marginal likelihood, also called model evidence. Its integration over a high-dimensional space—or summation for discrete variables—often makes exact posterior computation infeasible. Variational methods replace this calculation with an optimization problem over simpler distributions. Their development is closely associated with approximate inference in probabilistic graphical models. (people.eecs.berkeley.edu)
In the standard formulation, a family is chosen and the approximation is
where is Kullback–Leibler divergence. Its direction matters: exchanging the two distributions generally produces a different approximation. The restriction to determines which posterior features the method can represent. (cs.columbia.edu)
Evidence lower bound
Because the posterior contains the unknown evidence, implementations usually maximize the evidence lower bound (ELBO):
Here denotes an expectation under . The identity
shows that maximizing the ELBO is equivalent to minimizing the stated divergence for a fixed model. It also establishes that the ELBO cannot exceed the log evidence. (proceedings.mlr.press)
When , the bound becomes
The first term measures expected fit to observations; the second penalizes departure from the prior distribution. In latent-variable learning, this decomposition supports simultaneous optimization of generative-model parameters and variational parameters. The negative ELBO can serve as the training loss function. (arxiv.org)
Approximation families
A common choice is the mean-field family,
which imposes independence between selected variables or blocks in the approximation. This assumption does not assert that the true posterior is independent. It simplifies expectations and optimization, but cannot reproduce dependence between separately factorized blocks. Mean-field methods connect probabilistic inference with earlier approximation techniques in statistical mechanics. (people.eecs.berkeley.edu)
Structured families retain selected dependencies. For example, a full-covariance multivariate normal distribution can represent linear correlations that a diagonal Gaussian cannot. Hierarchical variational families introduce additional variables governing the approximation, allowing dependencies and richer distributional shapes. Increased flexibility can improve representation of the target, but also makes computation and optimization more demanding. The inference model’s expressiveness is distinct from the expressiveness of the underlying probabilistic model. (jmlr.csail.mit.edu)
Optimization algorithms
Coordinate-ascent variational inference updates one factor at a time while holding the others fixed. For unrestricted factors in a mean-field approximation, the optimal update has the form
In suitable conjugate models, these updates have closed forms. Coordinate updates improve the bound, but the overall problem generally need not be convex, so initialization can affect the solution. (jmlr.org)
Stochastic variational inference uses randomly selected observations or minibatches to estimate updates, avoiding a complete pass through the dataset at every iteration. The 2013 formulation by Hoffman and colleagues used stochastic optimization to scale Bayesian topic models to collections containing millions of documents. It distinguishes local latent variables associated with observations from global quantities shared across the dataset. (jmlr.org)
Black-box variational inference estimates ELBO gradients using samples from the approximation, reducing the need for model-specific derivations. Score-function estimators provide broad applicability, although their variance can be substantial; variance-reduction techniques improve their usability. Automatic differentiation further supports general-purpose implementations. ADVI combines transformations of constrained variables, Gaussian approximations, and automated differentiation without requiring conjugacy. (proceedings.mlr.press)
Amortized inference and neural models
Amortized inference learns a shared mapping from observations to approximate-posterior parameters instead of independently optimizing every observation’s parameters. In a variational autoencoder, an encoder neural network produces , while a decoder specifies . Both are trained through a variational bound. (arxiv.org)
The reparameterization trick expresses a sample as a differentiable transformation of parameter-independent noise. For a diagonal Gaussian,
This permits gradients to pass through the sampling transformation and supports minibatch training with stochastic gradient methods. Sharing an inference network reduces repeated computation, but limits approximate posteriors to those that the network can produce. (arxiv.org)
Accuracy and limitations
Unlike Markov chain Monte Carlo, standard VI generally retains approximation error even after optimization converges. Accuracy depends on the chosen family and the quality of optimization. Reverse-KL mean-field approximations can underestimate posterior uncertainty and concentrate on one region of a multimodal target. A stable ELBO therefore does not establish that posterior variances, tail probabilities, or dependencies are accurate. Faster computation and useful predictions can coexist with imperfect uncertainty estimates. (cs.columbia.edu)