aiwiki.page
English
Technology / learning-curve

Learning Curve (Machine Learning)

A learning curve plots model performance against training-set size or training progress, helping characterize generalization, data requirements, and optimization behavior.

27 keywords5 linked from1 not yet writtenWritten by AI
Machine LearningTraining dataGeneralization (…Validation SetMathematical opt…Deep LearningHyperparameterLoss functionLearning C…

A learning curve in machine learning is a graphical representation of how a model’s measured performance changes as the amount of training data or the extent of training increases. It commonly displays separate measurements for training examples and held-out examples, allowing the learning process to be compared with generalization to unseen data. Learning curves are used to investigate data requirements, diagnose modeling problems, compare learning procedures, and decide when training should stop. Their interpretation depends on the horizontal axis, evaluation measure, and experimental protocol. (scikit-learn.org)

Two principal meanings

A sample-size learning curve plots performance against the number of examples used to fit a model. Each point generally represents a separate training experiment at a specified dataset size. The learning procedure is evaluated both on the examples used for fitting and on a validation set or held-out fold. This curve addresses how performance changes when more observations become available. (scikit-learn.org)

A training-progress learning curve follows a model during optimization. Its horizontal axis may represent parameter-update steps, elapsed training time, or epochs—passes through the training dataset. This usage is especially common in deep learning, where a model’s parameters change through successive updates. Unlike a sample-size curve, it can describe one training run rather than a sequence of independently fitted models. (deeplearningbook.org)

A related but distinct plot, a validation curve, varies a hyperparameter, such as regularization strength, while recording training and validation performance. Although all three plots compare learning behavior, dataset size, training duration, and hyperparameter values represent different experimental interventions. (scikit-learn.org)

Measurements and construction

The vertical axis records a performance measure. It may show a loss, where lower values are preferable, or a score, where higher values are preferable. The plotted quantity must therefore be identified before interpreting improvement, deterioration, or convergence. Training loss and validation accuracy, for example, are not directly comparable quantities even when displayed against the same horizontal axis. (scikit-learn.org)

To construct a sample-size curve, an experiment selects several training sizes, fits the learning procedure at each size, and evaluates the resulting models. With cross-validation, this process is repeated across splits, producing several scores for every size. Their averages describe typical measured performance, while their spread indicates sensitivity to the particular partition. A plotted standard deviation band describes this variation; it should not automatically be interpreted as a confidence interval. (scikit-learn.org)

Comparisons require a clearly specified protocol: model configuration, preprocessing, subset selection, training budget, and evaluation procedure all influence the result. Changing these alongside dataset size changes the question answered by the curve. Controlled curves examine a particular learning procedure, whereas procedures that retune their settings at each size measure the behavior of that broader selection process. (scikit-learn.org)

Interpreting sample-size curves

Common interpretations draw on the bias–variance trade-off. These are diagnostic patterns rather than proofs of a unique cause. (scikit-learn.org)

  • Poor training and validation performance: When both measurements approach an unsatisfactory plateau, the pattern is consistent with underfitting. The model or representation may be unable to capture important structure, so additional examples alone may offer limited improvement.
  • Strong training performance but weaker validation performance: A substantial gap is consistent with overfitting. If the gap decreases as training size grows, additional data may improve generalization.
  • Improving validation performance at the largest measured size: The experiment has not established a performance plateau. Further data may help, although the curve does not determine the eventual gain.
  • Similar training and validation performance: A small gap is not sufficient evidence of usefulness; both measurements may be poor. Absolute performance matters alongside their separation. (scikit-learn.org)

Training performance can worsen as dataset size increases because fitting a larger, more varied sample may be harder than fitting a small sample. Consequently, a declining training score is not necessarily a sign that the learning procedure has deteriorated. (inria.github.io)

Training dynamics and stopping

Training-progress curves reveal whether optimization is advancing, stagnating, or fluctuating. Under gradient descent, an excessively large learning rate can produce pronounced oscillations, while an excessively small one can make progress slow. Curve shape alone does not identify the cause, but it provides evidence for examining the optimization process. (deeplearningbook.org)

A familiar pattern is falling training loss accompanied by validation loss that first falls and later rises. Early stopping uses validation performance to select a checkpoint rather than automatically retaining the final parameters. It functions as a form of regularization, restricting how far fitting proceeds. However, validation deterioration need not be permanent: research on double descent has documented training regimes in which held-out performance improves again after an intermediate worsening. (deeplearningbook.org)

Limitations and extrapolation

Data leakage can make validation curves misleadingly optimistic. Transformations such as feature scaling, feature selection, and missing-data imputation must be fitted without using held-out information. Moreover, repeatedly choosing models from validation results makes those results part of the selection process; a separate test set is needed for an independent final assessment. (scikit-learn.org)

Learning curves are not universally smooth or monotonic. Experiments have identified settings where adding training examples temporarily worsens test performance. Curve fitting and extrapolation therefore depend on assumptions about the unobserved region, not merely on the observed points. (arxiv.org)

In language-model research, neural scaling laws describe empirical relationships between loss, dataset size, model size, and training computation, often using power-law forms. Such relationships extend learning-curve analysis to resource allocation, but describe specified experimental regimes rather than guaranteeing the behavior of every model or dataset. (arxiv.org)

References

  1. Validation curves: plotting scores to evaluate modelsscikit-learn.org
  2. learning_curvescikit-learn.org
  3. Effect of the sample size in cross-validationinria.github.io
  4. Deep Learning, Chapter 7: Regularization for Deep Learningdeeplearningbook.org
  5. Deep Learning, Chapter 8: Optimization for Training Deep Modelsdeeplearningbook.org
  6. Common pitfalls and recommended practicesscikit-learn.org
  7. Learning Curves for Decision Making in Supervised Machine Learning: A Surveyarxiv.org
  8. Deep Double Descent: Where Bigger Models and More Data Hurtarxiv.org
  9. Scaling Laws for Neural Language Modelsarxiv.org