The beta distribution is a family of continuous probability distributions describing a random variable between zero and one. It has two positive shape parameters, usually denoted and , and is written . Its bounded range and varied shapes make it useful for representing proportions and uncertain probabilities. In particular, it provides an analytically tractable model for uncertainty about the success probability of a binary experiment. (docs.scipy.org)
Definition and distribution function
For and , the probability density function is
and is zero outside the unit interval. The normalizing constant is the beta function, defined by the integral
where denotes the gamma function. Thus, the density integrates to one for every positive parameter pair. (itl.nist.gov)
The cumulative distribution function is
Here is the regularized incomplete beta function. General quantiles have no simple closed-form expression and are evaluated numerically. Although densities can diverge near either endpoint, the distribution remains continuous: neither zero nor one receives positive probability mass. (itl.nist.gov)
Shapes and parameter interpretation
The parameters control how probability is distributed across the interval:
- When , the distribution is the uniform distribution on .
- When , the density has a single interior peak.
- When , it is U-shaped, with density diverging toward both endpoints.
- When and , excluding the uniform case, it decreases across the interval; reversing these inequalities produces an increasing density.
Equal parameters give symmetry about . Exchanging the parameters reflects the density: if , then . (randomservices.org)
A useful alternative parameterization separates the mean from concentration:
so that and . For fixed , increasing reduces dispersion without changing the mean. This follows directly from the variance formula below and helps distinguish a distribution’s central location from its degree of concentration. (itl.nist.gov)
Moments and scaling
The expected value and variance are
When both shape parameters exceed one, the mode is
These formulas describe the standard distribution on the unit interval. (itl.nist.gov)
A beta variable can be transformed to any finite interval , with , by setting . Its density becomes
Consequently, its mean is , and its variance is . The interval limits are additional location and scale specifications, not substitutes for the two shape parameters. (docs.scipy.org)
Bayesian inference
In Bayesian inference, a beta prior distribution is a conjugate prior for the success probability in a Bernoulli or binomial model. Suppose trials are conditionally independent given , and successes are observed. The likelihood function is proportional to . Multiplication by a beta prior, using Bayes’ theorem, gives the posterior distribution
Successes therefore increment the first parameter, while failures increment the second. (mas.ncl.ac.uk)
The posterior mean is
For , this can be written as a weighted average of the prior mean and the observed success fraction:
This algebra explains the interpretation of as a prior concentration or effective-weight parameter. The shape parameters need not be integers or represent literal historical counts. For example, a prior updated with eight successes and two failures becomes , with posterior mean . (mas.ncl.ac.uk)
Related distributions and estimation
The beta distribution arises from gamma distributions. If and have statistical independence, shapes , and the same positive rate, then
This identity provides a construction for generating beta random variables. (randomservices.org)
It also describes order statistics: the th smallest observation among independent uniform variables on has distribution . Thus, beta distributions occur naturally in the sampling distributions of ranked observations, not only as models chosen for bounded data. (math.arizona.edu)
The Dirichlet distribution generalizes the beta family to vectors of positive proportions summing to one. With two components, its first component has a beta distribution, while the second equals one minus the first. (docs.scipy.org)
For observations strictly between zero and one, maximum likelihood estimation determines the shape parameters by solving nonlinear equations. Numerical procedures are generally required. If the interval bounds are also estimated freely, the likelihood can become unbounded as fitted endpoints approach extreme observations; fixed bounds and freely estimated bounds therefore lead to substantially different fitting problems. (itl.nist.gov)