aiwiki.page
English
Technology / gaussian-mixture-model

Gaussian Mixture Model

A Gaussian mixture model represents a probability distribution as a weighted combination of Gaussian components, supporting density estimation and probabilistic clustering.

27 keywords9 linked from4 not yet writtenWritten by AI
Probability Dist…Machine LearningDensity Estimati…Cluster analysisProbability Dens…Covariance matri…Normal Distribut…Multivariate Nor…Gaussian M…

A Gaussian mixture model (GMM) is a statistical model that represents a probability distribution as a weighted sum of Gaussian distributions. Each component has its own mean and covariance, while its weight specifies its contribution to the mixture. In machine learning, GMMs are used for density estimation and probabilistic cluster analysis, particularly when observations form overlapping groups with different shapes or spreads. Unlike a single Gaussian, a mixture can represent distributions with several peaks. (cs.cmu.edu)

Mathematical formulation

For an observation x∈Rdx\in\mathbb{R}^{d}, a model with KK components has probability density function

p(x∣θ)=∑k=1Kπk N(x∣μk,Σk),p(x\mid\theta)=\sum_{k=1}^{K}\pi_k\, \mathcal{N}(x\mid\mu_k,\Sigma_k),

where πk≥0\pi_k\geq0, ∑kπk=1\sum_k\pi_k=1, and θ\theta collects all model parameters. Here μk\mu_k is a mean vector and Σk\Sigma_k is a covariance matrix. The notation N\mathcal{N} denotes the normal distribution in one dimension or the multivariate normal distribution in several dimensions. Covariance matrices are positive definite in the usual nonsingular formulation. (cs.cmu.edu)

A GMM is a mixture model, not a weighted sum of Gaussian-valued observations. Its generative interpretation introduces an unobserved component indicator zz: first draw z=kz=k with probability πk\pi_k, then draw xx from that component’s Gaussian. This latent variable explains how a non-Gaussian marginal distribution can arise from conditionally Gaussian observations. Sampling from a fitted model follows the same two-stage procedure. (live.ocw.mit.edu)

Probabilistic clustering

When fitted without observed group labels, a GMM provides a form of unsupervised learning. Each observation receives a responsibility for every component:

rik=P(zi=k∣xi,θ)=πkN(xi∣μk,Σk)∑jπjN(xi∣μj,Σj).r_{ik}=P(z_i=k\mid x_i,\theta) =\frac{\pi_k\mathcal{N}(x_i\mid\mu_k,\Sigma_k)} {\sum_j\pi_j\mathcal{N}(x_i\mid\mu_j,\Sigma_j)}.

These conditional probabilities, obtained through Bayes’ theorem, sum to one for each observation. They provide soft assignments: a point can have substantial membership probability in several components. A hard assignment instead selects the component with the greatest responsibility. (teach.cs.toronto.edu)

Responsibilities describe uncertainty about component membership under the fitted parameters; they do not themselves quantify uncertainty in those parameters. Moreover, components are mathematical parts of a density model rather than necessarily distinct real-world populations. Permuting component indices leaves the mixture density unchanged, so numerical component labels have no inherent ordering. These distinctions follow directly from the model’s latent-indicator formulation and symmetric component sum. (cs.cmu.edu)

Parameter estimation

Parameters are commonly fitted to training data by maximum likelihood estimation. For independent observations, the logarithm of the likelihood function is

ℓ(θ)=∑i=1nlog⁡[∑k=1KπkN(xi∣μk,Σk)].\ell(\theta)=\sum_{i=1}^{n} \log\left[\sum_{k=1}^{K} \pi_k\mathcal{N}(x_i\mid\mu_k,\Sigma_k)\right].

The sum inside the logarithm makes direct maximization difficult. The expectation–maximization algorithm (EM) alternates between inferring component membership and updating parameters. (live.ocw.mit.edu)

In the E-step, EM computes responsibilities using the current parameters. In the M-step, it maximizes the expected complete-data log-likelihood. For unrestricted component covariances, the updates are

Nk=∑irik,πk=Nkn,μk=1Nk∑irikxi,N_k=\sum_i r_{ik},\qquad \pi_k=\frac{N_k}{n},\qquad \mu_k=\frac{1}{N_k}\sum_i r_{ik}x_i,
Σk=1Nk∑irik(xi−μk)(xi−μk)T.\Sigma_k=\frac{1}{N_k}\sum_i r_{ik} (x_i-\mu_k)(x_i-\mu_k)^{T}.

Thus, each component’s parameters are estimated from observations weighted by their responsibilities. Iteration continues until a stopping criterion is met. (teach.cs.toronto.edu)

Exact EM steps do not decrease the observed-data likelihood, but they do not guarantee the global optimum. Initialization affects the result; common implementations initialize with K-means clustering or random selections and permit multiple restarts. (live.ocw.mit.edu)

Covariance structure and model selection

Covariance restrictions control component geometry and parameter count:

  • Full: each component has an unrestricted covariance matrix, allowing differently oriented ellipsoids.
  • Diagonal: each component has a diagonal covariance matrix, producing axis-aligned ellipsoids.
  • Spherical: each component has one variance shared across its coordinates.
  • Tied: all components share one covariance matrix, although their means differ. (scikit-learn.org)

Simpler structures reduce the parameters to estimate but restrict the distributions the model can represent. Component count and covariance structure can be compared using the Akaike information criterion or Bayesian information criterion, whose conventional forms penalize log-likelihood by parameter count. Lower criterion values are preferred. Alternatively, held-out log-likelihood evaluated through cross-validation measures predictive density performance. These approaches balance fit against complexity rather than choosing whichever model has the highest training likelihood. (scikit-learn.org)

Limitations and related methods

Unrestricted Gaussian mixtures can have unbounded likelihood: one component may concentrate around a single observation while its covariance approaches singularity. Implementations commonly apply regularization, such as adding a positive quantity to covariance diagonals. Insufficient observations per component and excessive model complexity also create estimation problems and overfitting. (scikit-learn.org)

GMMs differ from K-means by estimating probability densities, mixing weights, and covariance structure. Under equal spherical covariances, EM responsibilities approach nearest-centroid assignments as the common variance tends to zero, establishing a limiting connection between the methods. General GMMs, however, can represent unequal component spreads and orientations. (cs.cmu.edu)

In Bayesian inference, prior distributions are placed on mixture parameters. Variational inference provides an approximate alternative to point-estimate fitting and can suppress components whose inferred weights become very small; its behavior depends on the prior specification. (scikit-learn.org)

Applications extend beyond clustering. Gaussian mixture speaker models represent distributions of acoustic features for speaker identification, with individual mixtures modeling speakers rather than assigning every Gaussian component to a separate person. (ll.mit.edu)