Ensemble learning is a family of machine learning techniques that combine several models, called base learners or estimators, into a single predictive system. Rather than relying on one fitted model, an ensemble aggregates outputs through voting, averaging, weighted addition, or a learned combination rule. Its purpose is often to improve generalization or robustness, although combining models does not guarantee better performance. Ensembles can contain instances of the same learning method or models produced by different methods. (scikit-learn.org)
Principles and statistical foundations
An ensemble’s effectiveness depends on both the predictive quality of its members and the diversity of their errors. If all members make nearly identical mistakes, combining them offers little advantage. Diversity can arise from different samples of training data, feature subsets, model structures, or randomized training procedures. Random-forest theory explicitly relates classification performance to individual tree strength and the dependence between trees. (stat.berkeley.edu)
Averaging can reduce variance because fluctuations in different predictions partially cancel. The reduction is limited by correlation between members: adding highly similar predictors produces diminishing benefits. This connects ensemble learning to the bias–variance tradeoff. Bagging primarily targets instability, whereas sequential methods can construct increasingly expressive predictors. These descriptions are tendencies rather than universal rules; the effects depend on the learning problem, model family, and aggregation procedure. (doi.org)
Base learners need not all be weak models. Fully grown decision trees can be useful members of averaging ensembles, while shallow trees are common in boosting. The relevant property is how each learner contributes to the combined predictor, not simply whether its individual accuracy is low or high. (doi.org)
Bagging and random forests
Bagging, short for bootstrap aggregating, trains multiple versions of a predictor on resampled datasets. In its classical form, bootstrap sampling draws observations with replacement from the original dataset. Numerical predictions are averaged; classification predictions can be combined by a plurality vote. Leo Breiman’s 1996 paper showed that bagging can improve unstable learning procedures, for which small changes in training data cause substantial changes in predictions. (doi.org)
Decision trees are particularly suitable because their fitted structures can change considerably when the data change. A random forest adds further randomization by restricting the candidate features considered at each split. This reduces dependence between trees while retaining useful predictive strength. Breiman’s 2001 formulation established a theoretical framework connecting these properties to the forest’s generalization error. (stat.berkeley.edu)
Bootstrap sampling also enables out-of-bag evaluation. An observation is assessed using trees whose bootstrap samples excluded it, providing an internal estimate of predictive error. This reuses the training sample without fitting a separate model for every evaluation split, although it does not eliminate the need for appropriate external evaluation. (stat.berkeley.edu)
Boosting
Boosting constructs an ensemble sequentially, with each new learner responding to information from the current ensemble. AdaBoost adjusts observation weights so that previously misclassified examples receive greater emphasis. It then combines learners through a weighted vote. Yoav Freund and Robert Schapire presented its foundational formulation in work published as an extended abstract in 1995 and as a journal article in 1997. (sciencedirect.com)
Gradient boosting expresses learning as stagewise improvement of an additive model. Each new learner approximates the negative gradient of a chosen loss function with respect to current predictions. For squared-error regression, this corresponds to fitting residuals. Jerome Friedman’s 2001 paper developed this approach for regression and classification, allowing different losses to be handled within a common framework. (doi.org)
Boosting is therefore closely related to mathematical optimization in function space. Its behavior depends on learning rate, learner complexity, and the number of stages. Shrinkage, subsampling, and early stopping can control model complexity. Unlike independently trained bagging members, successive boosting stages depend on earlier stages, limiting straightforward parallel training across the sequence. (scikit-learn.org)
Voting, averaging, and stacking
Voting combines classification outputs directly. Hard voting selects the class receiving the largest vote total; soft voting averages predicted class probabilities and selects the highest-scoring class. Weights can give some members greater influence. For numerical prediction, an analogous approach averages model outputs. These methods can combine conceptually different predictors without learning a separate combination model. (scikit-learn.org)
Stacking, or stacked generalization, instead trains a meta-model using base-model predictions as inputs. For example, logistic regression can combine classification scores from several learners. The meta-model can learn relationships among their outputs rather than applying a fixed voting rule. (scikit-learn.org)
Training this second level requires care. With cross-validation, out-of-fold predictions are generated for observations excluded from each corresponding base-model fit. These predictions train the meta-model, while base learners can subsequently be refitted on all available training data. Using in-sample predictions instead creates a high risk of overfitting and misleading evaluation through data leakage. (scikit-learn.org)
Neural ensembles and evaluation
Ensembles also combine neural networks. Independently trained networks can produce different predictions through randomized initialization and training. Deep ensembles average their predictive distributions and have been studied as a scalable approach to uncertainty estimation. Their published evaluation includes behavior under distribution shift, but uncertainty quality remains an empirical property rather than an automatic consequence of model multiplicity. (arxiv.org)
Ensemble evaluation distinguishes fitting, selecting settings, and measuring final performance. A test set used repeatedly to select ensemble weights or other hyperparameters can itself become overfit. Cross-validation supports model selection, while an untouched test set assesses the selected system. Time-dependent or grouped observations require splitting procedures that preserve their structure rather than assuming every observation is interchangeable. (scikit-learn.org)