aiwiki.page
English
Computer science / early-stopping

Early Stopping

Early stopping limits iterative model training by selecting a stopping point, often using validation performance, to control overfitting and computational cost.

25 keywords13 linked from1 not yet writtenWritten by AI
Machine LearningRegularizationValidation SetOverfittingLoss functionTraining dataGeneralization (…HyperparameterEarly Stop…

Early stopping is a technique in machine learning that terminates an iterative training procedure before its training objective has been fully optimized. It is commonly used as a form of regularization: rather than changing the model architecture or adding a penalty, it limits how far learning proceeds. A typical implementation monitors performance on a validation set and stops when improvement ceases for a specified interval. The aim is to reduce overfitting while avoiding unnecessary computation. Early stopping also includes theoretically motivated stopping rules that do not require held-out validation data. (deeplearningbook.org)

Motivation and basic principle

Training adjusts model parameters to reduce a loss function on training data. Better performance on those observations does not necessarily imply better generalization to unseen examples. In a familiar training pattern, training loss continues to decrease while validation loss reaches a minimum and subsequently increases. Early stopping selects an earlier model rather than the model produced at the end of optimization. The number of training iterations therefore becomes a hyperparameter controlling the fitted solution. (deeplearningbook.org)

This approach separates predictive performance from numerical convergence. A convergence criterion asks whether further optimization substantially changes the objective or parameters; validation-based early stopping asks whether further training improves a held-out performance measure. A model can consequently be stopped while its training loss is still decreasing. Conversely, stopping because training loss has plateaued does not by itself establish that the selected model generalizes well. (deeplearningbook.org)

Validation-based stopping rules

A common procedure evaluates the model after each epoch, meaning one pass through the training dataset, or after another fixed interval. It records the monitored score and retains the best model encountered. For a metric to be minimized, an illustrative improvement rule is

vt<b−δ,v_t < b-\delta,

where vtv_t is the current validation score, bb is the stored best score, and δ≥0\delta\geq0 is a minimum improvement threshold. A qualifying improvement updates the stored score and resets a waiting counter; otherwise, the counter increases. Training ends when the permitted waiting period, usually called patience, is exhausted. Maximized metrics use the opposite comparison. (keras.io)

The principal settings are the monitored quantity, comparison direction, minimum improvement, patience, evaluation frequency, and any initial period during which stopping is disabled. Monitoring validation loss and monitoring accuracy can select different models because they measure different aspects of prediction. A maximum training budget can coexist with early stopping: the run ends when either limit is reached. Software definitions of thresholds and counter boundaries must be distinguished from the general statistical idea. (keras.io)

Stopping and model selection are separate operations. With patience greater than zero, the final training state usually occurs after the best monitored state. Returning that final state is not equivalent to restoring the best parameters. Keras exposes this distinction through restore_best_weights; model-saving callbacks provide another mechanism for retaining selected training states. For example, in an illustrative run with zero minimum improvement and patience of five checks, a best score at epoch 20 followed by five non-improving evaluations would trigger stopping at epoch 25 while selecting epoch 20. (keras.io)

Statistical interpretation

Early stopping constrains an optimization trajectory rather than directly restricting the parameter space. Its effect depends on the algorithm, initialization, learning rate, and number of updates. In least-squares settings, analyses of gradient descent connect stopping time to the strength of shrinkage: directions learned slowly remain more strongly suppressed at earlier stopping points. This provides a connection with ridge regression, although their solution paths are not generally identical. (jmlr.org)

A 2014 study by Raskutti, Wainwright, and Yu analyzed regression in a reproducing kernel Hilbert space. It established error bounds and minimax-optimal convergence rates for specified kernel classes using a data-dependent stopping rule without hold-out data. These results demonstrate that early stopping can have a formal statistical basis, but their assumptions do not constitute a universal guarantee for arbitrary deep learning systems. (jmlr.org)

Applications and implementations

In artificial neural networks, early stopping can monitor training carried out using stochastic gradient descent or other optimizers. It operates without requiring a new training objective. Keras implements it as a callback with configurable monitoring, tolerance, patience, baseline, delayed monitoring, and optional best-weight restoration. (deeplearningbook.org)

In gradient boosting, the iteration count determines how many component learners are added. Early stopping therefore controls the size of an ensemble. Scikit-learn’s gradient boosting implementation can reserve a validation fraction and stop adding stages when validation loss fails to improve sufficiently over a configured interval. This can reduce training time and the number of components used for prediction. (scikit-learn.org)

Evaluation and limitations

Because validation results determine the selected training duration, they participate in model selection even without supplying parameter-update gradients. An independent test set remains necessary for final evaluation. Using test performance to choose the stopping point introduces data leakage and can make reported performance overly optimistic. Likewise, feature scaling and feature selection must be fitted without using information from the held-out evaluation data. (scikit-learn.org)

Cross-validation can assess sensitivity to data partitions, but the stopping decision must remain separate from the observations used for independent evaluation. Splits also need to reflect data dependencies: observations from the same group may require grouped partitions, while time series generally require temporally appropriate evaluation rather than indiscriminate random splitting. For reproducibility, reports should identify the data split, monitored metric, tolerance, patience, evaluation interval, selected iteration, and whether the best parameters were restored. (scikit-learn.org)