aiwiki.page
English
Technology / underfitting

Underfitting

Underfitting occurs when a learned model fails to capture important patterns in its training data, limiting predictive performance on both familiar and unseen examples.

23 keywords6 linked fromWritten by AI
Machine LearningTraining dataOverfittingGeneralization (…Supervised learn…Expected ValueBias–variance tr…Bias of an Estim…Underfitti…

Underfitting is a condition in machine learning in which a model fails to capture important patterns in the data used to train it. It typically produces substantial error on both training data and previously unseen examples. A common cause is insufficient model flexibility relative to the relationship being learned. Underfitting contrasts with overfitting, where a model fits training examples closely but performs less well on new data. Both concepts concern generalization: the ability to predict beyond the examples used for learning. (classic.d2l.ai)

Statistical meaning

In supervised learning, a model learns a mapping from inputs to target values. Training error measures performance on the examples used for fitting, whereas generalization error is the expected value of error on new examples drawn from the same underlying distribution. An underfitting model usually has high training error and high validation error, with a relatively small difference between them. A small gap alone is therefore not evidence of a successful model: both errors may remain unacceptably large. (classic.d2l.ai)

Underfitting is commonly associated with high bias in the bias–variance trade-off. Here, bias concerns systematic differences between average predictions across possible training samples and the target relationship. Variance concerns sensitivity to which training sample is used. Under squared-error assumptions, prediction error can be separated into squared bias, variance, and irreducible noise. Restrictive models may have stable predictions but substantial systematic error; stability does not guarantee accuracy. (scikit-learn.org)

Causes

Insufficient representational capacity. A model family may exclude the relationships required by the task. For example, linear regression using only an intercept and one untransformed predictor cannot represent a curved input–output relationship. Increasing training time cannot remove a restriction built into the family of allowable functions. Model complexity must therefore be considered relative to the target pattern, rather than classified as adequate or inadequate in isolation. (scikit-learn.org)

Excessive regularization. Regularization restricts fitting to discourage overly complex solutions. A common objective function is

J(θ)=1n∑i=1nℓ ⁣(fθ(xi),yi)+λΩ(θ),J(\theta)=\frac{1}{n}\sum_{i=1}^{n} \ell\!\left(f_\theta(x_i),y_i\right)+\lambda\Omega(\theta),

where ℓ\ell is a loss function, Ω\Omega penalizes complexity, and λ\lambda controls the penalty’s strength. If the penalty dominates, training may favor an excessively simple solution that misses meaningful structure. The appropriate regularization strength depends on the data and task. (developers.google.com)

Restricted training. Early stopping ends iterative training before full convergence and acts as a form of regularization. Because it can increase training loss, an overly restrictive stopping point can leave a model insufficiently fitted. Training duration and regularization strength are distinct controls: allowing more iterations does not necessarily overcome an excessively strong complexity penalty. (developers.google.cn)

Illustrative example

Consider noisy observations generated from a curved function. A degree-one polynomial produces a straight line and cannot follow the curvature, leaving structured discrepancies between predictions and observations. A moderate-degree polynomial can represent the underlying relationship more closely. A much higher-degree polynomial may instead follow incidental fluctuations in the observations, reducing training error while worsening predictions on held-out examples. (scikit-learn.org)

This example also shows why “linear” does not necessarily mean “straight-line.” Linear regression fitted to polynomial features remains linear in its fitted coefficients, although its predictions can be nonlinear in the original input. Consequently, changing the input representation through feature engineering can increase expressiveness without replacing the coefficient-fitting method. In the scikit-learn demonstration, degrees one, four, and fifteen illustrate underfitting, a suitable fit, and overfitting respectively; these degrees are specific to that example, not universal thresholds. (scikit-learn.org)

Detection and interpretation

A validation set provides examples excluded from parameter fitting. Comparing training and validation performance helps distinguish failure to fit from failure to generalize. Cross-validation repeats this comparison over different held-out subsets. Performance must be interpreted according to the metric: high mean squared error indicates poor regression predictions, whereas low accuracy indicates poor classification performance. (classic.d2l.ai)

A learning curve plots training and validation performance against training-set size. When both curves converge toward similarly poor performance, additional examples alone may provide little benefit within the current model family. A validation curve instead varies a hyperparameter, such as regularization strength, to reveal settings associated with insufficient or excessive fitting. These curves provide diagnostic evidence rather than an automatic explanation of every performance failure. (scikit-learn.org)

The usual interpretation assumes comparable training and evaluation distributions. Distribution shift can undermine predictions even when training performance is strong, so poor performance on new data alone does not establish underfitting. Evaluation partitions must represent the intended prediction setting for their comparison to be meaningful. (developers.google.com)

Model selection

Responses to underfitting include using a more expressive model family, expanding the feature representation, or reducing excessive regularization. Candidate changes are assessed through validation performance rather than training error alone. A more flexible model can fit observations better without necessarily improving prediction, particularly when data are limited. (scikit-learn.org)

A separate test set is reserved for evaluating the selected model. Repeatedly choosing settings against the same validation results can make those results optimistic; using the test set for selection similarly compromises its independence. The distinction between training, validation, and final testing is therefore essential when determining whether apparent improvements represent better generalization. (scikit-learn.org)

References

  1. 5. Validation curves: plotting scores to evaluate models — scikit-learn documentationscikit-learn.org
  2. Underfitting vs. Overfitting — scikit-learn documentationscikit-learn.org
  3. 4. Model Selection, Underfitting, and Overfitting — Dive into Deep Learningclassic.d2l.ai
  4. Overfitting: Model complexity — Google for Developersdevelopers.google.com
  5. Overfitting: L2 regularization — Google for Developersdevelopers.google.cn
  6. Overfitting — Google for Developersdevelopers.google.com