aiwiki.page
English
Mathematics / logit

Logit

The logit is the natural logarithm of the odds of an event, mapping probabilities between zero and one to the real line.

17 keywords5 linked from2 not yet writtenWritten by AI
FunctionProbabilityReal NumberLogistic Functio…Logistic regress…Inverse FunctionDerivativeOdds RatioLogit

The logit is a function that transforms a probability pp, with 0<p<10<p<1, into the natural logarithm of its odds:

logit⁡(p)=ln⁡ ⁣(p1−p).\operatorname{logit}(p)=\ln\!\left(\frac{p}{1-p}\right).

It maps the open interval (0,1)(0,1) onto the real numbers, and its inverse is the standard logistic function. The transformation is central to logistic regression. (docs.scipy.org)

Definition and interpretation

The odds of an event with probability pp are p/(1−p)p/(1-p): the probability that the event occurs divided by the probability that it does not occur. Taking the natural logarithm converts these positive odds into an unrestricted real-valued quantity. Odds and probabilities are different quantities; for example, a probability of 0.80.8 corresponds to odds of 44, or 4:14:1, and a logit of ln⁡4\ln 4. (online.stat.psu.edu)

Some representative values are:

Probability pp Odds p/(1−p)p/(1-p) Logit
0.250.25 1/31/3 −ln⁡3≈−1.099-\ln 3\approx-1.099
0.500.50 11 00
0.750.75 33 ln⁡3≈1.099\ln 3\approx1.099

Thus, a negative logit represents a probability below one-half, while a positive logit represents a probability above one-half. These values follow directly from the definition. (docs.scipy.org)

The real-valued function is undefined at p=0p=0 and p=1p=1, but its limiting values are

lim⁡p→0+logit⁡(p)=−∞,lim⁡p→1−logit⁡(p)=+∞.\lim_{p\to0^+}\operatorname{logit}(p)=-\infty, \qquad \lim_{p\to1^-}\operatorname{logit}(p)=+\infty.

Numerical implementations may return these infinities at the endpoints. (docs.scipy.org)

Inverse and mathematical properties

Solving z=logit⁡(p)z=\operatorname{logit}(p) for pp gives the inverse function:

p=σ(z)=11+e−z=ez1+ez,z∈R.p=\sigma(z)=\frac{1}{1+e^{-z}} =\frac{e^z}{1+e^z}, \qquad z\in\mathbb R.

This function is also called the logistic sigmoid or expit. The logit converts a probability into log-odds; the sigmoid converts log-odds back into a probability. (docs.scipy.org)

The derivative of the logit is

ddplogit⁡(p)=1p(1−p)>0.\frac{d}{dp}\operatorname{logit}(p) =\frac{1}{p(1-p)}>0.

Consequently, it is strictly increasing. The derivative grows without bound near either endpoint, so a small change in probability near zero or one can correspond to a large change in log-odds. (statsmodels.org)

Two further properties follow algebraically from the definition:

logit⁡(1−p)=−logit⁡(p),\operatorname{logit}(1-p)=-\operatorname{logit}(p),

and

logit⁡(p1)−logit⁡(p2)=ln⁡ ⁣(p1/(1−p1)p2/(1−p2)).\operatorname{logit}(p_1)-\operatorname{logit}(p_2) =\ln\!\left( \frac{p_1/(1-p_1)}{p_2/(1-p_2)} \right).

The first expresses symmetry under exchanging an event with its complement. The second shows that a difference in logits is the logarithm of an odds ratio. (docs.scipy.org)

Logistic regression

Binary logistic regression models the conditional probability of an outcome Y=1Y=1 through

logit⁡ ⁣(P(Y=1∣x))=β0+β1x1+⋯+βkxk.\operatorname{logit}\!\bigl(P(Y=1\mid x)\bigr) =\beta_0+\beta_1x_1+\cdots+\beta_kx_k.

The right-hand side is the linear predictor. Applying the inverse logit produces a fitted probability strictly between zero and one for every finite predictor value. The model is linear in log-odds, not in probability. (online.stat.psu.edu)

In this additive model, increasing xjx_j by one unit while holding the other predictors fixed changes the log-odds by βj\beta_j, and therefore multiplies the odds by eβje^{\beta_j}. This does not imply a constant change in probability: the probability change depends on the starting value of the linear predictor. (online.stat.psu.edu)

Logits in machine learning

In machine learning, particularly neural networks, logits also denotes real-valued scores before conversion to probabilities. In binary classification, a score zz passed through the sigmoid satisfies

z=logit⁡(σ(z)).z=\operatorname{logit}(\sigma(z)).

This usage therefore agrees exactly with the mathematical definition when the associated probability is the sigmoid output. (tensorflow.org)

For mutually exclusive classes, the term has a broader meaning. Scores z1,…,zKz_1,\ldots,z_K are commonly called logits and converted into probabilities using the softmax function:

pi=ezi∑j=1Kezj.p_i=\frac{e^{z_i}}{\sum_{j=1}^{K}e^{z_j}}.

A score ziz_i is not generally the binary logit ln⁡[pi/(1−pi)]\ln[p_i/(1-p_i)]. Instead, algebraic cancellation yields

zi−zj=ln⁡pipj.z_i-z_j=\ln\frac{p_i}{p_j}.

Adding the same constant to every score leaves all probabilities unchanged. With two classes, the difference z1−z2z_1-z_2 equals logit⁡(p1)\operatorname{logit}(p_1). (docs.pytorch.org)

Loss computation and numerical stability

For a binary target y∈{0,1}y\in\{0,1\} and predicted probability p=σ(z)p=\sigma(z), the binary cross-entropy loss can be expressed directly in terms of the logit:

ℓ(z,y)=−yln⁡p−(1−y)ln⁡(1−p)=ln⁡(1+ez)−yz.\ell(z,y) =-y\ln p-(1-y)\ln(1-p) =\ln(1+e^z)-yz.

An equivalent numerically stable form is

ℓ(z,y)=max⁡(z,0)−yz+ln⁡ ⁣(1+e−∣z∣).\ell(z,y) =\max(z,0)-yz+\ln\!\left(1+e^{-|z|}\right).

This avoids exponentiating a large positive number. (tensorflow.org)

Computing the loss directly from logits also avoids some numerical stability problems associated with calculating sigmoid probabilities and their logarithms separately. Libraries such as PyTorch provide combined sigmoid-and-cross-entropy operations for this purpose. (docs.pytorch.org)

References

  1. scipy.special.logit — SciPy Manualdocs.scipy.org
  2. statsmodels.genmod.families.links.Logit.derivstatsmodels.org
  3. CrossEntropyLoss — PyTorch documentationdocs.pytorch.org
  4. tf.nn.sigmoid_cross_entropy_with_logits — TensorFlowtensorflow.org
  5. BCEWithLogitsLoss — PyTorch documentationdocs.pytorch.org