Missing data imputation is the process of replacing unavailable observations with values derived from observed information or an explicit model. In statistics and machine learning, it enables analyses and algorithms to operate on incomplete datasets. Missing entries may arise from unanswered questions, measurement failures, or information never recorded. Imputed values are estimates or plausible replacements, not recovered observations; their interpretation depends on the method and assumptions used. (scikit-learn.org)
Missingness mechanisms
The relationship between missingness and the data is central to choosing and interpreting an imputation procedure. Three conventional mechanisms describe this relationship:
- Missing completely at random (MCAR): the probability of missingness depends on neither observed nor unobserved data values.
- Missing at random (MAR): conditional on observed information, missingness does not additionally depend on the missing values.
- Missing not at random (MNAR): missingness still depends on unobserved values after conditioning on observed information.
MAR does not mean that missing entries occur haphazardly. For example, response rates may differ across observed groups without violating MAR. Rubin’s 1976 formulation established conditions under which likelihood-based or Bayesian inference can ignore the missingness process, including MAR and appropriate separation of model parameters. (people.csail.mit.edu)
Observed data alone generally cannot distinguish MAR from MNAR. An imputation method assuming MAR therefore does not establish that assumption’s truth. MNAR analysis requires additional assumptions or information; sensitivity analysis examines how results change under alternative assumptions about unobserved values. (journals.sagepub.com)
Single imputation methods
Single imputation produces one completed dataset. Simple methods replace missing entries with a constant, the observed mean or median for numerical variables, or the most frequent category for categorical variables. These methods are computationally inexpensive but do not model relationships among variables. (scikit-learn.org)
Multivariate methods exploit information across features. Regression imputation predicts an incomplete variable using other variables: linear regression can model numerical responses, while logistic regression can model binary responses. An iterative procedure cycles through incomplete variables, updating their replacements using the other variables’ observed or currently imputed values. Its predictive model may instead use a random forest or another suitable estimator. (search.r-project.org)
Neighbor-based imputation applies the k-nearest neighbors algorithm to identify similar records using available features. Missing numerical values are then replaced by averages, potentially distance-weighted, of observed values among selected neighbors. The definition of similarity and the availability of overlapping observed features affect the result. (scikit-learn.org)
Predictive mean matching combines modeling with donor selection. It identifies observed cases whose predicted values are close to a prediction for an incomplete case and draws an observed value from this donor group. Because replacements come from observed data, the method can preserve features of the variable’s empirical distribution that a simple parametric prediction may not reproduce. (stefvanbuuren.name)
Multiple imputation and uncertainty
Multiple imputation generates several plausible completed datasets rather than treating one replacement as certain. Each dataset is analyzed separately, and estimates are combined. Proper procedures represent uncertainty about missing values and imputation-model parameters; merely repeating a deterministic fill operation does not provide that uncertainty. (jstatsoft.org)
For a scalar parameter, let be its estimate and its estimated sampling variance in completed dataset , with datasets. Rubin’s rules give
Here, measures within-imputation uncertainty, measures between-imputation variance, and is the total variance used to calculate a standard error and appropriate confidence interval. Pooling combines analysis estimates, not the completed datasets into one averaged dataset. (stefvanbuuren.name)
Chained equations and model specification
Multiple imputation by chained equations (MICE), also called fully conditional specification, assigns a conditional model to each incomplete variable. Starting from initial replacements, it repeatedly updates variables using the other variables as predictors. Different models accommodate continuous, binary, ordered, and unordered categorical data. (jstatsoft.org)
Iterative imputation and multiple imputation are distinct: iteration describes how replacements are updated, whereas multiplicity describes generating several completed datasets. An iterative algorithm can return only one completion. (scikit-learn.org)
Model specification must reflect the intended analysis. Important interactions, nonlinear relationships, and multilevel structure may need explicit representation. Derived variables require consistent handling: independently imputing a total and its component variables can violate their arithmetic relationship. Passive imputation maintains such relationships by recalculating derived quantities from their components. (jstatsoft.org)
Predictive workflows and evaluation
For predictive modeling, imputation is a learned preprocessing step. Its parameters are estimated from training data, then applied without refitting to the validation set or test set. During cross-validation, fitting occurs separately within each training fold. Computing replacements before splitting can create data leakage and overly optimistic performance estimates. (scikit-learn.org)
Missingness indicators preserve information about which entries were absent. Some tree-based learners handle missing values directly, so explicit imputation is not always necessary. Reconstruction quality and downstream predictive performance are separate objectives; more elaborate imputation does not invariably improve prediction. (scikit-learn.org)
Evaluation can simulate missingness in data with known values and examine reconstruction error, estimator bias, uncertainty estimates, and interval coverage. Results depend on both the sampling process and the simulated missingness mechanism. Performance under one masking scheme therefore does not establish validity under every real-world missingness process. (stefvanbuuren.name)