A Gaussian process is a stochastic process in which every finite collection of indexed random variables follows a multivariate normal distribution, allowing degenerate distributions. Its index may represent time, spatial position, or the input to a mathematical model. A Gaussian process is completely specified by its mean function and covariance function. In statistics and machine learning, it provides a way to describe uncertainty about an unknown function rather than only about a finite set of parameters. (gaussianprocess.org)
Mathematical definition
Let be an index set and a collection of real-valued random variables. The collection is Gaussian if, for every positive integer and every choice , the vector
has a multivariate normal distribution. Requiring each individual value to have a normal distribution is insufficient: the joint distributions must also be Gaussian. (gaussianprocess.org)
The mean and covariance functions are
Here denotes expected value, while describes covariance between indexed values. The notation
implies that every finite evaluation vector has mean entries and covariance matrix entries . The diagonal value is the variance at . (gaussianprocess.org)
A valid covariance kernel is symmetric and positive semidefinite: every such finite matrix satisfies . These conditions ensure consistent Gaussian finite-dimensional distributions and permit construction of a process. The definition itself does not require continuous or differentiable sample paths; those properties depend on the kernel and additional regularity conditions. (gaussianprocess.org)
Covariance kernels and examples
The covariance kernel encodes assumptions about variation, smoothness, periodicity, and dependence across inputs. An important example is the squared-exponential, or radial basis function kernel,
Its amplitude controls marginal variance, and its length scale controls how quickly dependence decays with separation. It produces a process with mean-square derivatives of every order. Matérn kernels offer a parameter controlling finite mean-square differentiability, making different roughness assumptions possible. (gaussianprocess.org)
For a constant mean and a kernel depending only on , the resulting Gaussian process is a stationary process: translating all inputs leaves its finite-dimensional distributions unchanged. Valid kernels can also be combined through sums and products. The Wiener process is a nonstationary Gaussian example with zero mean and covariance for nonnegative times. (gaussianprocess.org)
Regression and prediction
Gaussian process regression uses Bayesian inference to update a prior distribution over functions using training data. A standard observation model is
with independent noise, also independent of . For Gaussian noise, conditioning gives an exact Gaussian posterior distribution over the latent function. (gaussianprocess.org)
Let , , and , where is the identity matrix. For a test input , define . Assuming is invertible, the posterior mean and variance are
Joint predictions at multiple inputs follow analogous matrix formulas. (gaussianprocess.org)
These equations distinguish uncertainty about the latent function from observation noise. A future noisy observation has predictive variance . A credible interval derived from this distribution is conditional on the assumed kernel, noise model, and fixed parameter values; it is not a model-independent guarantee of accuracy. (gaussianprocess.org)
Hyperparameters and model selection
Kernel amplitudes, length scales, and noise levels are hyperparameters. They may be estimated by maximizing the marginal likelihood, which integrates out latent function values. With , its logarithm is
The quadratic term measures data fit, while the determinant term penalizes the volume of outcomes permitted by the covariance model. Their balance supports model selection without reducing the criterion to training error alone. (gaussianprocess.org)
Optimization can have multiple local optima, and different parameter combinations can explain limited observations similarly. Fully Bayesian treatment instead assigns priors to hyperparameters and integrates over their posterior uncertainty. Even with Gaussian observation noise, this integration generally produces a mixture of conditional Gaussian predictions rather than a single Gaussian process posterior. (gaussianprocess.org)
Classification and related models
For classification, a latent Gaussian process is mapped to class probabilities, for example through a logistic function. The resulting non-Gaussian likelihood generally prevents exact Gaussian conditioning. Methods include Laplace approximation, expectation propagation, and Markov chain Monte Carlo. The latent prior remains Gaussian, but its posterior generally does not. (gaussianprocess.org)
Gaussian processes also connect to linear regression: Gaussian priors on linear-model weights induce Gaussian distributions over function evaluations. Kernel formulations extend this perspective to richer feature spaces without explicitly constructing every feature, linking Gaussian process prediction to kernel methods. (gaussianprocess.org)
Computational requirements
For observations, conventional exact regression with a dense covariance matrix requires storage and approximately factorization time. Implementations commonly use Cholesky decomposition and linear solves rather than explicitly forming an inverse. These costs limit straightforward applications to large datasets. (gaussianprocess.org)
Approximation methods include subsets of observations, reduced-rank covariance representations, inducing-variable methods, and iterative linear solvers. Some reduced-rank approaches use representative variables and have leading costs around . The computational savings depend on the construction and may alter predictive variances as well as means; approximation quality therefore concerns uncertainty estimates, not merely point predictions. (gaussianprocess.org)