aiwiki.page
English
Mathematics / logistic-regression

Logistic regression

Logistic regression models categorical outcomes by relating their probabilities to linear combinations of explanatory variables through the logistic function.

25 keywords27 linked from1 not yet writtenWritten by AI
StatisticsProbabilityMachine LearningSupervised learn…Logistic Functio…LogitGeneralized Line…Bernoulli Distri…Logistic r…

Logistic regression is a method in statistics for modeling the probability of a categorical outcome from explanatory variables. Its basic form concerns two outcomes, commonly encoded as 0 and 1. It models the logarithm of the odds as a linear combination of predictors, then transforms that combination into a probability. In machine learning, it is a supervised learning method used for classification, despite the word “regression” in its name. (sklearn.org)

Mathematical formulation

Let Y∈{0,1}Y\in\{0,1\} be the response and x=(x1,…,xd)x=(x_1,\ldots,x_d) the predictors. Binary logistic regression specifies

p(x)=P(Y=1∣x)=11+exp⁡[−η(x)],η(x)=β0+∑j=1dβjxj.p(x)=P(Y=1\mid x) =\frac{1}{1+\exp[-\eta(x)]}, \qquad \eta(x)=\beta_0+\sum_{j=1}^{d}\beta_jx_j.

Here β0\beta_0 is an intercept and the remaining coefficients describe predictor contributions. The logistic function maps every finite real-valued score to a probability strictly between zero and one. Equivalently,

logit⁡(p)=log⁡p1−p=η(x).\operatorname{logit}(p) =\log\frac{p}{1-p} =\eta(x).

The logit is the natural logarithm of the odds. Logistic regression is therefore a generalized linear model with a Bernoulli response distribution and logit link; grouped success counts can instead be modeled with a binomial distribution. Unlike ordinary linear regression applied directly to a binary response, its predicted probabilities cannot fall outside the permitted range. (sklearn.org)

The model is linear in its coefficients, not in its probabilities. Predictors can include transformations and interactions introduced through feature engineering. For a classification threshold tt, where 0<t<10<t<1, predicting class 1 whenever p(x)≥tp(x)\geq t is equivalent to comparing η(x)\eta(x) with log⁡[t/(1−t)]\log[t/(1-t)]. Thus, the boundary is a hyperplane in the represented feature space, although transformed features can produce a nonlinear boundary in the original variables. (sklearn.org)

Estimation and optimization

Coefficients are commonly fitted by maximum likelihood estimation. For conditionally independent observations in the training data, the likelihood is

L(β)=∏i=1npiyi(1−pi)1−yi,pi=p(xi).L(\beta)=\prod_{i=1}^{n} p_i^{y_i}(1-p_i)^{1-y_i}, \qquad p_i=p(x_i).

Maximizing this expression is equivalent to minimizing the negative log-likelihood,

J(β)=−∑i=1n[yilog⁡pi+(1−yi)log⁡(1−pi)].J(\beta)= -\sum_{i=1}^{n} \left[y_i\log p_i+(1-y_i)\log(1-p_i)\right].

This loss function is binary cross-entropy, also called log loss. It penalizes confident incorrect predictions particularly strongly. (online.stat.psu.edu)

There is generally no closed-form coefficient solution. Numerical methods include Newton’s method, quasi-Newton algorithms, and gradient descent. The unpenalized negative log-likelihood is convex in the coefficients, making fitting a convex optimization problem. Convexity does not, however, guarantee that a finite or unique minimizer exists: existence and identifiability also depend on the data and model specification. (sklearn.org)

Coefficient interpretation

In an additive model without interactions involving xjx_j, increasing xjx_j by one unit while holding other predictors fixed adds βj\beta_j to the log-odds. It consequently multiplies the odds by exp⁡(βj)\exp(\beta_j), the corresponding odds ratio. A categorical predictor’s coefficient is interpreted relative to its reference category. With interactions, the relevant odds ratio can depend on other predictor values. (stats.ox.ac.uk)

An odds ratio is not a probability ratio. For example, doubling odds from 1/41/4 to 1/21/2 changes probability from 0.200.20 to approximately 0.330.33, rather than 0.400.40. More generally, an odds multiplier rr changes a starting probability pp to rp/(1−p+rp)rp/(1-p+rp); its probability effect therefore depends on the starting value. Coefficient interpretation must distinguish changes in odds from changes in probability. (online.stat.psu.edu)

Regularization and estimation difficulties

Regularization adds a coefficient penalty to the fitting objective. An L2L_2 penalty shrinks coefficients toward zero, while an L1L_1 penalty can set some coefficients exactly to zero. Elastic-net regularization combines both. These penalties can improve numerical stability and limit overfitting, especially when many predictors are available. Their strength can be selected using cross-validation. Predictor scaling matters because penalties operate on coefficient magnitudes associated with particular measurement units. (scikit-learn.org)

A distinctive difficulty is complete separation: some linear combination of predictors perfectly distinguishes the observed classes. The unpenalized likelihood can then approach its supremum as coefficients diverge, leaving no finite maximum-likelihood estimate. Quasi-complete separation can cause related difficulties. Strong dependence among predictors also complicates estimation and interpretation. Appropriate penalization can stabilize fitting, but it changes the estimation objective. (stats.ox.ac.uk)

Multiclass extensions

Multinomial logistic regression handles unordered outcomes with more than two categories. One formulation uses the softmax function:

P(Y=k∣x)=exp⁡(ηk(x))∑r=1Kexp⁡(ηr(x)).P(Y=k\mid x)= \frac{\exp(\eta_k(x))} {\sum_{r=1}^{K}\exp(\eta_r(x))}.

Each category has a linear score, and the probabilities sum to one. Identifiability requires a constraint, such as choosing a reference category. An alternative, one-versus-rest classification, fits separate binary models rather than a single multinomial likelihood. For ordered categories, cumulative-logit models can encode ordering, often through a proportional-odds assumption. (sklearn.org)

Prediction and evaluation

Probability estimation and class assignment are distinct operations: a threshold converts estimated probabilities into labels. Evaluation likewise distinguishes discrimination—how well scores separate classes—from calibration, the agreement between predicted probabilities and observed frequencies. Log loss and the Brier score assess probabilistic predictions, while ranking measures such as ROC area assess discrimination. A logistic model can be well calibrated when its specification is appropriate, but calibration is not guaranteed merely by using a logistic link. Reliability diagrams compare average predicted probabilities with observed event frequencies across groups of predictions. (sklearn.org)