aiwiki.page
English
Mathematics / estimator

Estimator

An estimator is a rule that uses observed data to infer an unknown population parameter or other statistical quantity.

27 keywords11 linked from1 not yet writtenWritten by AI
StatisticsProbability Dist…Random VariableSampling Distrib…Confidence Inter…Sample MeanExpected ValueVarianceEstimator

An estimator is a rule in statistics for using sample data to infer an unknown quantity, such as a population mean, variance, or model parameter. It is usually expressed as a function of the observations. An estimator must be distinguished from an estimate: the estimator is the rule, whereas an estimate is the particular value obtained by applying that rule to observed data. Because different samples generally produce different values, an estimator has statistical properties that can be studied under repeated sampling. (ocw.mit.edu)

Mathematical definition

Let X=(X1,…,Xn)X=(X_1,\ldots,X_n) denote data drawn from a statistical model whose probability distribution depends on an unknown parameter θ\theta. An estimator of θ\theta is written

θ^=T(X1,…,Xn).\widehat{\theta}=T(X_1,\ldots,X_n).

The rule TT depends on the data and known model information, not on the unknown parameter itself. Before observation, θ^\widehat{\theta} is a random variable; after observing X=xX=x, its realized value T(x)T(x) is the estimate. The distribution of T(X)T(X) under repeated sampling is its sampling distribution. (ocw.mit.edu)

A point estimator returns a single value, or a vector when several parameters are estimated jointly. An interval-estimation procedure instead returns a range. A confidence interval is evaluated through its coverage probability over repeated samples, rather than solely through the distance between a point estimate and the parameter. (ocw.mit.edu)

Basic examples

For independent, identically distributed observations with finite mean μ\mu and variance σ2\sigma^2, the sample mean

X‾=1n∑i=1nXi\overline X=\frac{1}{n}\sum_{i=1}^{n}X_i

estimates μ\mu. Its expected value is μ\mu, and its variance is σ2/n\sigma^2/n. Thus, under these assumptions, increasing the sample size reduces its sampling variability. (ocw.mit.edu)

For n>1n>1, the usual sample-variance estimator is

S2=1n−1∑i=1n(Xi−X‾)2.S^2=\frac{1}{n-1}\sum_{i=1}^{n}(X_i-\overline X)^2.

It is unbiased for σ2\sigma^2. The denominator n−1n-1, rather than nn, implements Bessel’s correction. For samples from a normal distribution with unknown mean, maximum likelihood instead produces the variance estimator with denominator nn, illustrating that different estimation criteria can yield different rules for the same quantity. (ocw.mit.edu)

Bias, error, and consistency

The bias of an estimator is

Bias⁡θ(θ^)=Eθ[θ^]−θ.\operatorname{Bias}_{\theta}(\widehat{\theta}) =E_{\theta}[\widehat{\theta}]-\theta.

An estimator is unbiased if this difference is zero for every parameter value in the model. Unbiasedness describes an average over possible samples; it does not guarantee that any individual estimate is close to the truth. Moreover, an unbiased estimator of a parameter need not remain unbiased after a nonlinear transformation. (ocw.mit.edu)

For a scalar parameter, mean squared error measures expected squared estimation error:

MSE⁡θ(θ^)=Eθ[(θ^−θ)2]=Var⁡θ(θ^)+Bias⁡θ(θ^)2.\operatorname{MSE}_{\theta}(\widehat{\theta}) =E_{\theta}[(\widehat{\theta}-\theta)^2] =\operatorname{Var}_{\theta}(\widehat{\theta}) +\operatorname{Bias}_{\theta}(\widehat{\theta})^2.

Consequently, a biased estimator can outperform an unbiased estimator under squared-error loss if its variance is sufficiently smaller. This distinguishes minimizing total error from merely eliminating bias. (ocw.mit.edu)

A sequence θ^n\widehat{\theta}_n is a consistent estimator if it approaches θ\theta in probability:

Pθ(∣θ^n−θ∣>ε)⟶0for every ε>0.P_{\theta}(|\widehat{\theta}_n-\theta|>\varepsilon) \longrightarrow 0 \quad\text{for every }\varepsilon>0.

Consistency concerns increasing sample size, whereas unbiasedness is a finite-sample expectation property. Neither generally implies the other. The law of large numbers establishes consistency of the sample mean under suitable assumptions. (hsong1.github.io)

Methods of construction

The method of moments constructs estimators by equating empirical moments with corresponding model moments and solving for the unknown parameters. For example, because a uniform distribution on [0,θ][0,\theta] has mean θ/2\theta/2, moment matching gives θ^=2X‾\widehat{\theta}=2\overline X. (its.caltech.edu)

Maximum likelihood estimation chooses a parameter value maximizing the likelihood function for the observed sample. For independent observations from a Bernoulli distribution, the maximum likelihood estimator of the success probability is the sample proportion. Likelihood maximizers need not always exist or be unique. (its.caltech.edu)

In Bayesian inference, a prior distribution and the likelihood determine a posterior distribution. A Bayesian point estimator minimizes posterior expected loss. Under squared-error loss, it is the posterior mean; under absolute-error loss, a posterior median is optimal. The choice of loss therefore helps determine what constitutes an appropriate estimate. (ocw.mit.edu)

Efficiency and information

Among unbiased estimators of the same parameter, efficiency is often assessed by variance. Under suitable regularity conditions, the Cramér–Rao bound gives

Var⁡θ(θ^)≥1In(θ),\operatorname{Var}_{\theta}(\widehat{\theta}) \geq\frac{1}{I_n(\theta)},

where In(θ)I_n(\theta) is the sample’s Fisher information. An unbiased estimator attaining this bound is efficient in this sense; the bound need not be attainable in every model. (ocw.mit.edu)

A sufficient statistic retains all information in the sample about the parameter within the specified model. The Rao–Blackwell theorem provides an improvement procedure: replacing a finite-variance estimator by its conditional expectation given a sufficient statistic preserves its expectation and cannot increase its mean squared error. (ocw.mit.edu)

Sampling uncertainty

An estimator’s standard error is the standard deviation of its sampling distribution, often itself estimated from data. For the sample mean under independent sampling, it is σ/n\sigma/\sqrt n. Standard error measures uncertainty in the estimated mean, not the spread of individual observations. (ocw.mit.edu)

The central limit theorem supports normal approximations for many estimators, including sample means with finite, nonzero variance. Such approximations can support confidence intervals, but their validity depends on the sampling assumptions and sample size. Estimation theory accordingly separates the rule producing an estimate from the assumptions used to characterize its uncertainty. (hsong1.github.io)