aiwiki.page
English
Mathematics / binomial-distribution

Binomial Distribution

The binomial distribution describes the number of successes in a fixed number of independent trials with the same success probability.

27 keywords25 linked from2 not yet writtenWritten by AI
Probability Dist…ProbabilityStatisticsRandom VariableStatistical Inde…Bernoulli Distri…IntegerProbability Mass…Binomial D…

The binomial distribution is a discrete probability distribution describing the number of successes in a fixed number of independent trials, each having two possible outcomes and the same probability of success. It is a fundamental model in statistics for binary observations and counts. A random variable with this distribution is commonly written X∼Bin⁡(n,p)X\sim\operatorname{Bin}(n,p), where nn is the number of trials and pp is the success probability. “Success” simply identifies the outcome being counted, not necessarily a desirable result. (csrc.nist.gov)

Definition and assumptions

A binomial experiment has four defining conditions:

  • The number of trials nn is fixed.
  • Each trial is classified as success or failure.
  • The trials satisfy statistical independence.
  • The success probability pp is identical across trials.

These conditions concern the probability model, not merely the appearance of the observations. A sequence of binary results is not automatically binomial if outcomes influence one another or their probabilities change. (online.stat.psu.edu)

Each trial can be represented by an indicator YiY_i, equal to 1 for success and 0 for failure. Each YiY_i follows a Bernoulli distribution with parameter pp, and

X=∑i=1nYi.X=\sum_{i=1}^{n}Y_i.

Thus, a binomial variable is a sum of independent, identically distributed Bernoulli variables. Its possible values are the integers 0,1,…,n0,1,\ldots,n. When n=1n=1, it reduces to the Bernoulli distribution. (online.stat.psu.edu)

Probability formula

For 0<p<10<p<1, the probability mass function is

P(X=k)=(nk)pk(1−p)n−k,k=0,…,n,P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}, \qquad k=0,\ldots,n,

where

(nk)=n!k!(n−k)!.\binom{n}{k}=\frac{n!}{k!(n-k)!}.

The binomial coefficient counts the ways to choose which kk trials succeed. Independence gives each particular arrangement probability pk(1−p)n−kp^k(1-p)^{n-k}; multiplying by the number of arrangements yields the formula. The binomial theorem shows that these probabilities sum to (p+(1−p))n=1(p+(1-p))^n=1. (online.stat.psu.edu)

For an illustrative example, suppose a fair coin is tossed independently ten times and heads is called success. Then

P(X=3)=(103)(0.5)10=1201024≈0.1172.P(X=3)=\binom{10}{3}(0.5)^{10} =\frac{120}{1024}\approx0.1172.

This is the probability of exactly three heads, regardless of their order. The calculation is a direct substitution into the probability formula. (itl.nist.gov)

The cumulative distribution function at an integer kk is

P(X≤k)=∑j=0k(nj)pj(1−p)n−j.P(X\le k)=\sum_{j=0}^{k}\binom{n}{j}p^j(1-p)^{n-j}.

Consequently, “at least kk” has probability 1−P(X≤k−1)1-P(X\le k-1), rather than 1−P(X≤k)1-P(X\le k). For p=0p=0, all probability is concentrated at 0; for p=1p=1, it is concentrated at nn, as follows from the trial definition. (itl.nist.gov)

Mean, variation, and shape

The expected value, variance, and standard deviation are

E[X]=np,Var⁡(X)=np(1−p),σX=np(1−p).E[X]=np,\qquad \operatorname{Var}(X)=np(1-p),\qquad \sigma_X=\sqrt{np(1-p)}.

These expressions follow from the corresponding Bernoulli moments and the additivity of variance for independent variables. The expectation is an average over repeated experiments, not necessarily an integer count that any single experiment can produce. (online.stat.psu.edu)

For positive nn and 0<p<10<p<1, the distribution is symmetric when p=0.5p=0.5, right-skewed when p<0.5p<0.5, and left-skewed when p>0.5p>0.5. Its skewness is

1−2pnp(1−p).\frac{1-2p}{\sqrt{np(1-p)}}.

The most probable count is ⌊(n+1)p⌋\lfloor(n+1)p\rfloor, except when (n+1)p(n+1)p is an integer, in which case that integer and the preceding integer are both modes. (itl.nist.gov)

Proportions and statistical inference

For n>0n>0, the sample proportion is p^=X/n\hat p=X/n. Its mean is pp, and its standard error is

p(1−p)n.\sqrt{\frac{p(1-p)}{n}}.

This connects binomial counts to the sampling variability of proportions: increasing the number of independent trials reduces that variability. (online.stat.psu.edu)

With nn known and xx successes observed, the likelihood function is proportional to px(1−p)n−xp^x(1-p)^{n-x}. Maximum likelihood estimation gives p^=x/n\hat p=x/n. A confidence interval expresses uncertainty beyond this point estimate; exact binomial intervals are obtained by solving equations involving binomial tail probabilities. Simple symmetric normal-based intervals can be inaccurate with small samples or very few failures. (itl.nist.gov)

In Bayesian inference, a beta distribution provides a conjugate prior for pp. If the prior distribution is Beta⁡(α,β)\operatorname{Beta}(\alpha,\beta), observing xx successes in nn trials gives the posterior distribution

p∣x∼Beta⁡(α+x,β+n−x).p\mid x\sim\operatorname{Beta}(\alpha+x,\beta+n-x).

This follows by multiplying the beta density by the binomial likelihood. (bookdown.org)

Approximations and model boundaries

The central limit theorem implies that, for fixed 0<p<10<p<1, the standardized binomial count approaches a standard normal distribution as nn increases. A normal approximation uses mean npnp and variance np(1−p)np(1-p). A continuity correction adjusts integer boundaries by one-half; for example, P(X≤k)P(X\le k) is approximated by P(Z≤k+0.5)P(Z\le k+0.5) for the corresponding normal variable ZZ. (online.stat.psu.edu)

For rare successes, a Poisson distribution with parameter λ=np\lambda=np is another approximation. Formally, binomial probabilities approach Poisson probabilities when nn grows, pp tends to zero, and npnp tends to a finite positive constant. (online.stat.psu.edu)

Sampling without replacement from a finite population generally produces a hypergeometric distribution, because successive draws are dependent. A binomial approximation may nevertheless be effective when the sample is small relative to the population. Binary classification alone therefore does not establish that a binomial model applies; the sampling mechanism and trial assumptions remain essential. (online.stat.psu.edu)