aiwiki.page
English
Mathematics / nested-cross-validation

Nested Cross-validation

Nested cross-validation separates model selection from performance estimation by placing an inner validation procedure inside an outer evaluation loop.

29 keywords7 linked from1 not yet writtenWritten by AI
Machine LearningStatisticsCross-validationGeneralization (…OverfittingHyperparameterRegularizationFeature selectio…Nested Cro…

Nested cross-validation is a model-evaluation procedure in machine learning and statistics that uses two levels of cross-validation to separate model selection from performance assessment. The inner level selects a model configuration, usually through hyperparameter optimization, while the outer level estimates the generalization performance of the complete selection-and-training procedure. Its central purpose is to avoid evaluating a selected configuration on the same validation results that led to its selection. (scikit-learn.org)

Motivation and selection bias

Ordinary cross-validation estimates performance by repeatedly fitting a model on one subset and evaluating it on another. However, when many configurations are compared, the configuration with the best validation score may benefit partly from random variation rather than genuinely superior predictive performance. Reporting that winning score as an independent performance estimate can therefore produce optimistic results. This is overfitting of the selection criterion, distinct from fitting model parameters too closely to the training observations. (jmlr.org)

Selection can concern a hyperparameter, such as regularization strength, or broader choices such as model family and feature selection. A study comparing support vector machines, random forests, and logistic regression may place all these alternatives inside the inner selection procedure. If their outer scores instead determine the winner, the reported winning outer score becomes subject to another layer of selection bias. Nesting separates selection from evaluation; it does not make unrestricted reuse of evaluation results harmless. (jmlr.org)

The two-level procedure

In a conventional implementation, the dataset is partitioned into KK outer folds. Each fold serves once as an outer test set, while the remaining observations form the outer training data. Within each outer training set, a separate JJ-fold cross-validation creates inner training subsets and a rotating validation set. The outer test fold remains excluded throughout selection. (scikit-learn.org)

For each outer fold, the procedure is:

  1. Evaluate candidate configurations using only the inner folds.
  2. Select the configuration with the best aggregate inner score.
  3. Refit that configuration on the entire outer training set.
  4. Evaluate the refitted model once on the outer test fold.

The outer scores are then aggregated. Different outer folds can select different configurations; the object being evaluated is the selection procedure, not necessarily one fixed configuration. (scikit-learn.org)

Mathematical formulation

For supervised learning, let D={(xi,yi)}i=1nD=\{(x_i,y_i)\}_{i=1}^{n}, and let IkI_k contain the indices of outer test fold kk. Write D−kD_{-k} for the remaining data. Let AA denote the full learning procedure, including inner selection and refitting, and define

f^k=A(D−k).\widehat f_k=A(D_{-k}).

For a pointwise loss function ℓ\ell, a sample-weighted nested estimate is

R^nested=1n∑k=1K∑i∈Ikℓ ⁣(yi,f^k(xi)).\widehat R_{\mathrm{nested}} =\frac{1}{n}\sum_{k=1}^{K} \sum_{i\in I_k} \ell\!\left(y_i,\widehat f_k(x_i)\right).

This expresses the outer evaluation of the complete procedure. With equally sized folds, it equals the arithmetic mean of the fold-average losses. With unequal folds, weighting by fold size gives each observation equal weight. For nonlinear metrics, such as the F-score, averaging fold scores and scoring pooled predictions need not yield the same quantity. (scikit-learn.org)

The estimate concerns a procedure trained on the outer training-set size, approximately n(K−1)/Kn(K-1)/K, rather than directly measuring a final model trained on all nn observations. Consequently, separation of tuning and testing does not guarantee exact absence of estimator bias for every target of interest. (sklearn.org)

Preprocessing and data leakage

Nesting is effective only when every data-dependent operation respects the split boundaries. Feature scaling, missing-data imputation, and principal component analysis must be fitted using the relevant training subset, then applied to validation or test observations without refitting on them. Selecting features using the complete dataset before cross-validation introduces data leakage, even if model fitting itself follows a nested design. (scikit-learn.org)

A processing pipeline keeps transformations and estimation together so that each inner training fold fits its own preprocessing steps. After selection, the entire pipeline is refitted on the outer training set. Outer test observations must not determine transformation parameters, feature choices, or other learned components. (scikit-learn.org)

The splitting strategy also determines the meaning of evaluation. Random folds suit settings where observations are sufficiently independent and exchangeable. Group-aware splits keep related observations together when performance on unseen groups is the target. For time series, chronological splits can evaluate prediction of later observations from earlier ones. These constraints apply at both nesting levels. (sklearn.org)

Computational cost and uncertainty

With KK outer folds, JJ inner folds, and MM candidate configurations, exhaustive selection entails approximately KJMKJM candidate fits plus KK outer refits. This count follows directly from repeating the inner search for each outer fold. Parallel computing can reduce elapsed time, but does not remove the underlying fitting workload. (scikit-learn.org)

Outer scores show variation across partitions, but they are not independent replicates because their training sets overlap. Their standard deviation is therefore not automatically a valid standard error of the aggregate estimate. A confidence interval calculated as though fold scores were independent can understate uncertainty. Bengio and Grandvalet established that no universally unbiased variance estimator exists for ordinary KK-fold cross-validation. (jmlr.org)

Final fitting and reporting

After evaluation, the same predefined selection procedure can be run on the complete development dataset, followed by refitting the selected configuration on all available development observations. The nested scores remain estimates of the procedure’s performance, not fresh measurements of that final fitted model. A separately reserved test set provides another assessment only if it remains untouched by development decisions. (scikit-learn.org)

For reproducibility, an evaluation description records both splitting schemes, candidate configurations, scoring and aggregation rules, preprocessing steps, random seeds, and refitting behavior. These details specify which selection procedure the reported outer scores actually evaluate. (sklearn.org)