The Bernoulli distribution is a discrete probability distribution for a random variable that takes only the values 0 and 1. Its single parameter, , is the probability of observing 1; the probability of observing 0 is . It describes one binary observation, such as whether a coin lands heads or whether a component passes inspection. The labels “success” and “failure” conventionally denote 1 and 0 without implying that either outcome is desirable. (online.stat.psu.edu)
Definition and representation
Writing , with , the probability mass function is
For , this can be expressed compactly as
At , the variable equals 0 with certainty; at , it equals 1 with certainty. These are degenerate cases of the same family. (online.stat.psu.edu)
A Bernoulli variable can also represent the occurrence of an event . Its indicator random variable, , equals 1 when occurs and 0 otherwise, and therefore has parameter . The underlying experiment need not have only two elementary outcomes: rolling a die and recording whether the result is six produces a Bernoulli variable, even though the die itself has six possible results. (probabilitycourse.com)
Mean, variance, and uncertainty
The expected value and variance are
Because , the second moment is also , giving the variance immediately as . Consequently, the standard deviation is . The variance is greatest at , where it equals , and vanishes at the two endpoints. Unlike many distribution families, its mean and variance cannot be selected independently. These properties follow directly from its mass function. (online.stat.psu.edu)
In information theory, its entropy, measured in bits, is
where is interpreted as 0 by continuity. Entropy is zero for a certain outcome and reaches one bit when both outcomes are equally probable. It measures uncertainty about the observation, rather than uncertainty about an estimated parameter. (web.stanford.edu)
Repeated observations and the binomial distribution
If are independent Bernoulli variables sharing the same parameter , their sum
has the binomial distribution with parameters and . Thus, a Bernoulli distribution is the special case of a binomial distribution with one trial. The distinction is between recording a single outcome and counting successes across several trials. (online.stat.psu.edu)
Both statistical independence and a common success probability matter for this result. Binary observations can individually be Bernoulli without their sum having the stated binomial distribution: observations may depend on one another, or their success probabilities may differ. A Bernoulli marginal distribution alone does not specify how observations are related. (online.stat.psu.edu)
For independent observations with common parameter , the sample proportion has mean and variance . The central limit theorem explains why its sampling distribution approaches a normal distribution as the sample size increases, provided . This concerns an aggregate, not the distribution of an individual binary observation. (online.stat.psu.edu)
Parameter estimation
In statistics, suppose independent observations contain successes. The likelihood function for the observed sequence is
Its logarithm is
Maximum likelihood estimation gives
the observed fraction of successes. When both outcomes occur, this follows by differentiating the log-likelihood and solving for its maximum. If every observation is 0 or every observation is 1, the maximum lies at the corresponding endpoint. (stat135.berkeley.edu)
In Bayesian inference, a beta distribution is a conjugate prior for . With prior distribution , observing successes and failures gives the posterior distribution
Here the Bernoulli distribution describes observations conditional on , whereas the beta distribution describes uncertainty about itself. (bayesball.github.io)
Binary prediction models
In machine learning, logistic regression models a binary response conditionally on explanatory variables:
The probability can therefore vary between observations while each conditional response remains Bernoulli. This is a binary-response generalized linear model. (web.stanford.edu)
For an observed label and predicted probability , the negative log-likelihood is
This is binary cross-entropy, used as a loss function. It evaluates probability predictions rather than only whether a thresholded classification is correct: assigning very low probability to the outcome that actually occurs produces a large loss. (cs229.stanford.edu)