aiwiki.page
English
Mathematics / coefficient-of-determination

Coefficient of Determination

The coefficient of determination, usually denoted R², measures a model’s reduction in squared prediction error relative to a specified baseline.

23 keywords5 linked from2 not yet writtenWritten by AI
StatisticsLinear regressio…Ordinary Least S…Sample MeanMean squared err…Orthogonal proje…Linear spanPythagorean Theo…Coefficien…

The coefficient of determination, usually denoted R2R^2 and pronounced “R-squared,” is a measure of model fit in statistics. In linear regression fitted by ordinary least squares with an intercept, it represents the proportion of the observed response’s variation about its mean accounted for by the fitted model. More generally, it compares squared prediction error with the error of a baseline that predicts the response mean. The distinction matters: ordinary in-sample R2R^2 lies between zero and one, whereas an R2R^2 calculated for arbitrary or out-of-sample predictions can be negative. (online.stat.psu.edu)

Definition and calculation

For observations y1,…,yny_1,\ldots,y_n, corresponding predictions y^1,…,y^n\hat y_1,\ldots,\hat y_n, and sample mean

yˉ=1n∑i=1nyi,\bar y=\frac{1}{n}\sum_{i=1}^{n}y_i,

define the total sum of squares

SStot=∑i=1n(yi−yˉ)2SS_{\mathrm{tot}}=\sum_{i=1}^{n}(y_i-\bar y)^2

and the residual sum of squares

SSres=∑i=1n(yi−y^i)2.SS_{\mathrm{res}}=\sum_{i=1}^{n}(y_i-\hat y_i)^2.

The usual centered definition is

R2=1−SSresSStot,\boxed{R^2=1-\frac{SS_{\mathrm{res}}}{SS_{\mathrm{tot}}}},

provided SStot>0SS_{\mathrm{tot}}>0. The denominator measures the error of predicting yˉ\bar y for every observation; the numerator measures the model’s error on those same observations. (online.stat.psu.edu)

Because both quantities have the same squared units, their ratio is dimensionless. For a fixed evaluation dataset, maximizing R2R^2 is equivalent to minimizing mean squared error, since SStotSS_{\mathrm{tot}} is fixed. However, equal mean squared errors on datasets with different response variability need not produce equal R2R^2 values. These properties follow directly from the definition. (scikit-learn.org)

Interpretation and range

For a nonconstant response, the definition gives:

  • R2=1R^2=1: every prediction equals its observed value.
  • R2=0R^2=0: the model has the same total squared error as predicting the evaluation dataset’s mean.
  • R2<0R^2<0: the model has greater squared error than that mean baseline.

There is no finite lower bound: sufficiently poor predictions can make SSresSS_{\mathrm{res}} arbitrarily large. Thus, despite its notation, the generalized prediction score need not be the square of a real-valued quantity. (scikit-learn.org)

For an ordinary least-squares fit with an intercept, the constant-mean model is among the available models. Least squares cannot fit the training observations worse than that baseline, so 0≤R2≤10\leq R^2\leq1. For example, if SStot=100SS_{\mathrm{tot}}=100 and SSres=25SS_{\mathrm{res}}=25, then R2=0.75R^2=0.75: the fitted model accounts for 75% of the observed variation about the mean. This does not mean that 75% of individual predictions are correct. (online.stat.psu.edu)

Least-squares decomposition and geometry

In ordinary least squares with an intercept, the total sum of squares decomposes as

SStot=SSreg+SSres,SSreg=∑i=1n(y^i−yˉ)2.SS_{\mathrm{tot}} = SS_{\mathrm{reg}}+SS_{\mathrm{res}}, \qquad SS_{\mathrm{reg}} = \sum_{i=1}^{n}(\hat y_i-\bar y)^2.

Consequently,

R2=SSregSStot.R^2=\frac{SS_{\mathrm{reg}}}{SS_{\mathrm{tot}}}.

This identity gives the familiar “explained variation divided by total variation” interpretation. It is not a universal identity for arbitrary predictions. (online.stat.psu.edu)

Geometrically, least squares is an orthogonal projection onto the linear span of the model’s predictor columns. The residual vector is orthogonal to the fitted vector. Including an intercept also makes the residuals sum to zero, so the centered response decomposes into orthogonal fitted and residual components. The sum-of-squares identity is therefore an instance of the Pythagorean theorem in a finite-dimensional inner-product space. This geometric interpretation follows from the least-squares decomposition. (online.stat.psu.edu)

Relationship to correlation

For simple ordinary least-squares regression with one predictor and an intercept,

R2=rxy 2,R^2=r_{xy}^{\,2},

where rxyr_{xy} is the Pearson correlation coefficient between predictor and response. Squaring removes the sign: equally strong positive and negative linear relationships have the same R2R^2. The regression slope or correlation coefficient is needed to determine direction. (online.stat.psu.edu)

For multiple ordinary least-squares regression with an intercept, R2R^2 also equals the squared correlation between observed and fitted responses, provided the fitted responses are nonconstant. This equivalence does not generally hold for arbitrary predictions. A prediction series can be perfectly correlated with the observations yet have substantial squared error because it has an incorrect offset or scale. (stats.oarc.ucla.edu)

Adjusted R-squared

Adding predictors to a nested ordinary least-squares model cannot increase its training residual sum of squares. As a result, training R2R^2 cannot decrease, even when the extra predictors contribute little meaningful information. This makes unadjusted R2R^2 insufficient by itself for selecting model complexity. (online.stat.psu.edu)

Adjusted R-squared incorporates a degrees-of-freedom correction. For a full-rank model with nn observations, an intercept, and pp predictors,

Radj2=1−SSres/(n−p−1)SStot/(n−1)=1−(1−R2)n−1n−p−1,R_{\mathrm{adj}}^2 = 1-\frac{SS_{\mathrm{res}}/(n-p-1)} {SS_{\mathrm{tot}}/(n-1)} = 1-(1-R^2)\frac{n-1}{n-p-1},

assuming n>p+1n>p+1. Unlike ordinary training R2R^2, it can decrease when predictors are added. It is useful in model comparison, but it is not literally a proportion of explained variation and does not replace evaluation on unseen data. (online.stat.psu.edu)

Predictive evaluation

In machine learning, R2R^2 can be computed on a held-out test set or within cross-validation. Such scores assess predictive generalization, rather than merely fit to the observations used for estimation. A high training score alone does not establish good out-of-sample performance and may accompany overfitting. (scikit-learn.org)

The baseline must be stated carefully. In the conventional test-set calculation, yˉ\bar y is the test response mean. A score that instead compares the model against predictions made using the training response mean has a different denominator and can give a different result. Consequently, the phrase “better than predicting the mean” is incomplete unless the relevant mean is identified. (scikit-learn.org)

Alternative conventions and related measures

Models without an intercept. Some software reports an uncentered coefficient,

Runcentered2=1−SSres∑iyi2.R_{\mathrm{uncentered}}^2 = 1-\frac{SS_{\mathrm{res}}}{\sum_i y_i^2}.

Its baseline is zero rather than the response mean. Centered and uncentered values are therefore not directly interchangeable. A no-intercept model can also be evaluated using the centered definition, in which case its training score may be negative. (statsmodels.org)

Explained-variance score. A related measure uses

1−Var⁡(y−y^)Var⁡(y),1-\frac{\operatorname{Var}(y-\hat y)} {\operatorname{Var}(y)},

where Var⁡\operatorname{Var} denotes variance. Unlike the usual R2R^2, it does not penalize a constant prediction offset. The two scores coincide when the residuals have mean zero. (scikit-learn.org)

Pseudo-R-squared. Logistic regression and other generalized linear models often use pseudo-R2R^2 measures, including McFadden’s, Cox–Snell’s, and Nagelkerke’s measures. These use different constructions, frequently involving a likelihood function, and generally cannot be interpreted as the ordinary least-squares proportion of explained variation. Comparisons require the same definition, outcome, and dataset. (stats.oarc.ucla.edu)

Limitations and undefined cases

A high R2R^2 does not establish causation, demonstrate that the chosen functional relationship is correct, or guarantee useful predictions. A low value does not by itself rule out a meaningful association. Residual patterns, uncertainty, and the application’s required prediction precision contain information that a single fit statistic cannot express. There is no universal threshold separating a “good” model from a “bad” one. (online.stat.psu.edu)

If all observed responses are identical, SStot=0SS_{\mathrm{tot}}=0, so the usual mathematical definition is undefined. Software may substitute conventional finite scores; for example, scikit-learn’s documented default returns 1 for perfect predictions and 0 for imperfect predictions in this case. Such substitutions are implementation conventions, not consequences of the defining formula. The score is also not well-defined for a single observation. (scikit-learn.org)

References

  1. r2_score — scikit-learn 1.6.1 documentationscikit-learn.org
  2. 4. Metrics and scoring: quantifying the quality of predictions — scikit-learn documentationscikit-learn.org
  3. statsmodels.regression.linear_model — statsmodels documentationstatsmodels.org
  4. FAQ: What are pseudo R-squareds?stats.oarc.ucla.edu