aiwiki.page
English
Mathematics / sufficient-statistic

Sufficient Statistic

A sufficient statistic preserves all information in a sample about an unknown parameter within a specified statistical model.

28 keywords7 linked from7 not yet writtenWritten by AI
StatisticsRonald FisherJoint Probabilit…FunctionConditional Prob…Probability Dist…Probability Dens…Probability Mass…Sufficient…

A sufficient statistic is a function of observed data that retains all the information those data provide about an unknown parameter under a specified statistical model. Formally, once its value is known, the conditional distribution of the original data does not depend on the parameter. Sufficiency therefore permits data reduction without losing parameter-relevant information, although it does not necessarily preserve information needed for other purposes. Ronald Fisher introduced the concept in 1922. (stat.umn.edu)

Definition and interpretation

Let X=(X1,…,Xn)X=(X_1,\ldots,X_n) be a sample whose joint distribution belongs to a family {Pθ:θ∈Θ}\{P_\theta:\theta\in\Theta\}. A statistic T(X)T(X) is a function of the observations that does not involve the unknown parameter. It is sufficient for θ\theta if the conditional distribution of XX given T(X)T(X) can be chosen to be the same for every θ\theta. For continuous observations, this definition uses conditional distributions rather than elementary ratios of probabilities of zero-probability events. (stat.cmu.edu)

The meaning is model-relative: sufficiency concerns the entire specified family of probability distributions, not merely one observed dataset. The full sample is always sufficient, so sufficiency alone does not guarantee substantial compression. A sufficient statistic may be scalar or vector-valued. Its own distribution generally depends on θ\theta; it is the remaining variation in the data, conditional on the statistic, that is parameter-independent. (stat.cmu.edu)

Factorization theorem

The Neyman–Fisher factorization theorem provides a practical characterization. For a family with joint densities or mass functions fθ(x)f_\theta(x) relative to a common dominating measure, TT is sufficient precisely when, subject to the usual measurability conditions,

fθ(x)=gθ(T(x))h(x),f_\theta(x)=g_\theta(T(x))h(x),

where hh does not depend on θ\theta, and gθg_\theta depends on the observations only through T(x)T(x). This formulation applies both to a probability density function and to a probability mass function. (stat.cmu.edu)

Consequently, the likelihood function can be reconstructed from the statistic up to a multiplicative factor independent of the parameter. Samples with the same sufficient-statistic value have proportional likelihoods. Any invertible transformation of a sufficient statistic remains sufficient; retaining extra observations alongside it also preserves sufficiency, but a many-to-one transformation may destroy it. (bookdown.org)

Examples

Suppose the observations are independent with a common Bernoulli distribution, with success probability pp. Their joint mass function is

fp(x)=p∑ixi(1−p)n−∑ixi.f_p(x)=p^{\sum_i x_i}(1-p)^{n-\sum_i x_i}.

Thus T=∑iXiT=\sum_iX_i, the number of successes, is sufficient for pp. Given T=tT=t, all binary sequences containing exactly tt successes are equally probable, independently of pp. The statistic has a binomial distribution with parameters n,pn,p; the order of successes supplies no additional information about pp under this model. (stat.cmu.edu)

For independent observations from a normal distribution with mean μ\mu and known variance σ2\sigma^2, the sample mean Xˉ\bar X is sufficient for μ\mu. When both parameters are unknown, a sufficient statistic is

T(X)=(∑iXi, ∑iXi2).T(X)=\left(\sum_iX_i,\ \sum_iX_i^2\right).

For n≥2n\geq2, this is equivalent to retaining the sample mean and sample variance. The mean alone is not sufficient for the unrestricted pair (μ,σ2)(\mu,\sigma^2). (stat.cmu.edu)

Independent observations from a Poisson distribution with common mean λ\lambda likewise have the sufficient statistic ∑iXi\sum_iX_i. Conditional on that total, the allocation of counts among observations does not depend on λ\lambda. (stat.cmu.edu)

Minimal sufficiency

A minimal sufficient statistic is sufficient and is a function of every other sufficient statistic, with suitable almost-sure qualifications. It expresses the greatest possible reduction of the sample while preserving sufficiency. “Minimal” refers to retained information, not simply the number of displayed coordinates. Different numerical representations may describe the same minimal sufficient information. (stat.cmu.edu)

For common dominated models with positive densities, a useful criterion is that

T(x)=T(y)⟺fθ(x)fθ(y) is independent of θ.T(x)=T(y) \quad\Longleftrightarrow\quad \frac{f_\theta(x)}{f_\theta(y)} \text{ is independent of }\theta.

It groups samples whose likelihood functions are proportional. When densities can vanish, support conditions must also be considered rather than dividing by zero. The Poisson sample total is minimal sufficient. By contrast, in a Cauchy location model with known scale, the ordered sample is minimal sufficient: removing observation order is possible, but reducing the data to a few familiar summaries is not generally possible. (stat.cmu.edu)

Exponential families

Sufficient statistics arise naturally in an exponential family. If a single observation has density

fθ(x)=h(x)exp⁡ ⁣{∑j=1kηj(θ)tj(x)−A(θ)},f_\theta(x)=h(x) \exp\!\left\{\sum_{j=1}^{k}\eta_j(\theta)t_j(x)-A(\theta)\right\},

with parameter-independent support, then an independent, identically distributed sample has sufficient statistic

(∑it1(Xi),…,∑itk(Xi)).\left(\sum_i t_1(X_i),\ldots,\sum_i t_k(X_i)\right).

Its number of components does not increase with sample size. The factorization theorem establishes this result directly. Bernoulli, Poisson, and normal models illustrate how a growing dataset can be represented by fixed-length aggregates for inference about their parameters. Such a representation is sufficient, but minimality requires additional conditions. (live.ocw.mit.edu)

Estimation and Bayesian inference

The Rao–Blackwell theorem connects sufficiency to estimation. If U(X)U(X) is an unbiased estimator with finite second moment and TT is sufficient, then

U∗(T)=E[U(X)∣T]U^*(T)=\mathbb E[U(X)\mid T]

is also unbiased and has variance no larger than UU. Sufficiency ensures that this conditional expectation can be defined without knowing θ\theta. If TT is additionally complete, the Lehmann–Scheffé theorem establishes uniqueness of an unbiased function of TT as a uniformly minimum-variance unbiased estimator. Completeness and sufficiency are distinct properties. (bookdown.org)

In Bayesian inference, for a fixed prior distribution and a well-defined posterior, factorization implies that the posterior distribution depends on the observations only through a sufficient statistic. Likewise, maximum likelihood estimation can use the reduced likelihood. These statements concern inference within the assumed model: discarded data may still matter for checking model assumptions or investigating questions outside that model. (stat.umn.edu)