Density estimation is the task in statistics and machine learning of estimating an unknown probability density function from observations. Rather than describing data only through averages or variances, it reconstructs the distribution’s shape, including concentrations, asymmetry, and multiple peaks. Methods range from fitting a specified family of distributions to constructing flexible estimates whose complexity grows with the available data. It is commonly treated as unsupervised learning, because observations need not carry target labels. (stat.cmu.edu)
Mathematical formulation
Suppose are independent observations from the same continuous probability distribution, with unknown density . A density estimator is a data-dependent function intended to approximate . A valid estimated density satisfies
For a random variable , the estimated probability of an interval is obtained through an integral:
Density height is therefore not itself an event probability; probability corresponds to area under the curve. (stat.cmu.edu)
Estimating a density differs from estimating a cumulative distribution function. The empirical cumulative distribution assigns equal probability mass to each observation and is a step function. Although it estimates cumulative probabilities directly, it does not provide a smooth density between observed values. Density estimation introduces a model or smoothing mechanism to represent those intervening regions. (stat.cmu.edu)
Parametric and mixture methods
Parametric estimation assumes that belongs to a family described by finitely many parameters. For example, a normal distribution is characterized by its mean and variance. Parameters can be fitted using maximum likelihood estimation, which maximizes
This approach can be statistically efficient when the assumed family is appropriate, but its shape restrictions remain even with large samples. A single normal density, for instance, cannot represent two separated peaks. (stat.cmu.edu)
A Gaussian mixture model provides greater flexibility:
Each component has its own mean and covariance, while the weights determine its contribution. Parameters are commonly fitted with the expectation–maximization algorithm. Selecting the number of components is a separate model-selection problem. Unconstrained Gaussian-mixture likelihoods can become unbounded when a component collapses around an observation; covariance regularization limits this degeneracy. (scikit-learn.org)
Histograms and kernel estimators
A histogram partitions the observation space into bins. In one dimension, a bin containing observations and having width receives density . The resulting estimate is piecewise constant. Both bin width and bin placement affect its appearance: shifting boundaries can change apparent peaks despite leaving the data unchanged. (stat.cmu.edu)
Kernel density estimation replaces bins with localized contributions centered on observations:
Here is a nonnegative kernel integrating to one, and is the bandwidth. A Gaussian kernel contributes a bell-shaped bump around each observation; averaging the bumps produces a smooth density. Other choices include uniform and Epanechnikov kernels. Using Gaussian bumps does not imply that the underlying distribution is Gaussian. (stat.cmu.edu)
Bandwidth controls the bias–variance trade-off. Small bandwidths preserve fine detail but amplify sampling fluctuations and may produce spurious peaks. Large bandwidths reduce variability but can obscure genuine structure. Bandwidth usually matters more than the precise kernel shape. Selection procedures include reference-distribution rules, plug-in estimates, and cross-validation. (stat.cmu.edu)
Accuracy and dimensionality
A common theoretical criterion is mean integrated squared error:
It combines squared estimator bias and sampling variance across the domain. Under suitable regularity conditions, a one-dimensional kernel estimate is a consistent estimator when its bandwidth decreases while increases without bound. For sufficiently smooth densities and conventional second-order kernels, the asymptotically optimal bandwidth scales as , giving MISE of order . These rates depend on assumptions about smoothness and the estimator. (stat.cmu.edu)
Predictive performance can also be assessed through held-out log likelihood, closely related to Kullback–Leibler divergence. Evaluation on observations excluded from fitting helps distinguish generalizable distributional structure from overfitting. (stat.cmu.edu)
Multivariate estimation targets a joint probability distribution. Kernel methods extend to several dimensions, but local neighborhoods become sparsely populated as dimension increases—the curse of dimensionality. Boundary bias is another limitation: ordinary kernels can allocate mass outside the permitted domain, distorting estimates near its edges. (scikit-learn.org)
Neural and conditional density estimation
Flexible neural approaches include autoregressive models and normalizing flows. A flow transforms a simple base distribution through invertible, differentiable mappings. The change-of-variables formula, involving a Jacobian matrix determinant, gives the transformed density explicitly. Architecture determines the computational trade-offs between fitting, density evaluation, and sampling. (jmlr.org)
Conditional density estimation models , rather than an unconditional density. It represents the entire distribution of an outcome given explanatory variables, including changing spread and multiple possible modes, rather than predicting only a conditional mean. Such models support probabilistic prediction, while unconditional estimates also serve distribution visualization and generative sampling. (stat.cmu.edu)
References
- 36-402, Undergraduate Advanced Data Analysis (2011)stat.cmu.edu
- Estimating Distributions and Densitiesstat.cmu.edu
- 8. Density Estimation — scikit-learn 1.5.2 documentationscikit-learn.org
- 1. Gaussian mixture models — scikit-learn 1.1.3 documentationscikit-learn.org
- Supervised Learningstat.cmu.edu
- Data Visualizationstat.cmu.edu
- Visualizing Quantitative Distributionsstat.cmu.edu
- All of Nonparametric Statisticsstat.cmu.edu
- Normalizing Flows for Probabilistic Modeling and Inferencejmlr.org