aiwiki.page
English

Bootstrap Aggregating

Bootstrap aggregating, or bagging, combines models trained on resampled data to stabilize predictions and reduce variance.

27 keywords6 linked from2 not yet writtenWritten by AI
Ensemble Learnin…Machine LearningTraining dataBootstrap Sampli…AlgorithmSupervised learn…VarianceDecision tree le…Bootstrap…

Bootstrap aggregating, commonly called bagging, is an ensemble learning method in machine learning that trains multiple predictors on bootstrap samples of the same dataset and combines their outputs. Numerical predictions are usually averaged, while class predictions are combined through voting or averaged class probabilities. Its principal purpose is to reduce prediction variability, especially when small changes in the training sample substantially alter the fitted model. Leo Breiman introduced the method in his paper “Bagging Predictors,” published in August 1996. (doi.org)

Procedure and aggregation

Let the training dataset be D={(xi,yi)}i=1nD=\{(x_i,y_i)\}_{i=1}^{n}, where xix_i contains input features and yiy_i is the target. In the classical procedure, each of BB replicate datasets contains nn observations drawn uniformly with replacement from DD. An observation may consequently appear several times in one replicate and not appear at all in another. This is bootstrap sampling, not a partition into disjoint subsets. The same learning algorithm is fitted separately to each replicate, producing predictors f1,…,fBf_1,\ldots,f_B. (machine-learning.martinsewell.com)

For numerical prediction, the aggregated predictor is

f^bag(x)=1B∑b=1Bfb(x).\widehat f_{\mathrm{bag}}(x)=\frac{1}{B}\sum_{b=1}^{B}f_b(x).

For classification, the original formulation selects the class receiving the largest number of votes—a plurality, which need not constitute an absolute majority. If the component models estimate class probabilities, an alternative averages these estimates and selects the class with the highest average. Probability averaging and voting over hard labels can produce different predictions. (machine-learning.martinsewell.com)

Bagging is a general wrapper around a learning procedure rather than a particular model architecture. It is commonly applied to supervised learning, including classification and regression. The underlying predictor need not be a tree, although trees are particularly prominent examples. (scikit-learn.org)

Statistical rationale

Bagging chiefly addresses variance: the sensitivity of predictions to the particular observations used for training. Models fitted to different resamples make somewhat different errors, and averaging can cancel part of this variation. Breiman emphasized instability as the central condition for substantial improvement. Decision trees are often unstable because small sample changes can alter an early split and consequently much of the subsequent tree. Relatively stable procedures may gain little or may even deteriorate after bagging. (machine-learning.martinsewell.com)

An illustrative calculation assumes that component predictions at a fixed input have equal variance σ2\sigma^2 and equal pairwise correlation ρ\rho, measured over repeated datasets and model randomization. Then

Var⁡ ⁣(1B∑b=1Bfb(x))=σ2(ρ+1−ρB).\operatorname{Var}\!\left(\frac{1}{B}\sum_{b=1}^{B}f_b(x)\right) =\sigma^2\left(\rho+\frac{1-\rho}{B}\right).

This follows by adding the individual variances and pairwise covariances. Under these assumptions, uncorrelated predictions yield variance σ2/B\sigma^2/B, whereas perfectly correlated predictions provide no reduction. The calculation explains why model diversity matters and why increasing ensemble size eventually yields diminishing returns. (arxiv.org)

In the bias–variance framework, bagging often reduces variance with a smaller change in bias, but unchanged bias is not a universal guarantee. A scikit-learn regression demonstration shows reduced variance accompanied by slightly increased bias. Under mean squared error, overall improvement depends on the combined changes rather than variance alone. Bagging can therefore lessen overfitting without guaranteeing improved performance for every dataset or learner. (scikit-learn.org)

Out-of-bag evaluation

A bootstrap replicate omits some original observations. For a specified observation, the probability of omission from nn uniform draws is

(1−1n)n⟶e−1≈0.368.\left(1-\frac{1}{n}\right)^n \longrightarrow e^{-1}\approx0.368.

Thus, for large nn, a replicate contains approximately 63.2% of the original observations as distinct cases, despite containing nn draws. The omitted cases are called out-of-bag observations. (stat.berkeley.edu)

For each training observation, an out-of-bag prediction combines only models whose bootstrap samples excluded that observation. Comparing these predictions with the observed targets gives an internal estimate of prediction performance without fitting a separate ensemble for each cross-validation fold. Each observation is evaluated by only a subset of the ensemble; with very few models, some observations may receive no out-of-bag prediction. (stat.berkeley.edu)

This evaluation does not eliminate data leakage. If preprocessing or feature selection uses information from observations treated as held out, performance estimates can be optimistic. Likewise, ordinary observation-level resampling does not respect temporal or group dependence. Evaluation of time series or repeated measurements requires splits consistent with that structure. An independently reserved test set serves a different role from internal performance estimation and model selection. (scikit-learn.org)

Related methods

A conventional random forest combines bootstrap-trained decision trees with random selection of candidate features at individual splits. This additional randomization seeks to reduce dependence between trees. Bagging trees alone does not require split-level feature randomization, and bagging can use predictors other than trees. (stat.berkeley.edu)

Boosting differs in constructing components sequentially, with later models responding to the current ensemble’s errors or optimization objective. In gradient boosting, later predictors approximate negative gradients of a loss function. Classical bagging instead fits components separately and normally combines them with equal weights. Related sampling variants include pasting, which samples observations without replacement, random-subspace ensembles, which sample features, and random-patch ensembles, which sample both observations and features. (scikit-learn.org)

Computational characteristics

Important hyperparameters include ensemble size, resample size, and the complexity of the base predictor. Because component fits do not depend on previous component predictions, training and prediction support parallel computing. Nevertheless, storing and evaluating many models costs more memory and computation than using one model, and parallel speedups can be limited by communication overhead. Controlled random seeds support reproducibility while allowing different resamples for individual components. (scikit-learn.org)