aiwiki.page
English
Mathematics / statistical-hypothesis-testing

Statistical Hypothesis Testing

Statistical hypothesis testing evaluates claims about populations or models using sample data and procedures with specified error properties.

26 keywords44 linked from2 not yet writtenWritten by AI
StatisticsHypothesisNull HypothesisAlternative Hypo…Probability Dist…Random VariableTest StatisticSampling Distrib…Statistica…

Statistical hypothesis testing is a framework in statistics for evaluating claims about populations or data-generating processes using observed data. A test specifies competing hypotheses, a measure of discrepancy, and a rule for deciding whether the observations warrant rejecting a designated hypothesis. Its conclusions depend on the statistical model and sampling assumptions; they do not establish truth with certainty. (itl.nist.gov)

Hypotheses and statistical models

A statistical hypothesis specifies a property of a population or model, such as a mean, a proportion, or the relationship between variables. The null hypothesis, denoted H0H_0, is the claim tested. The alternative hypothesis, denoted H1H_1 or HaH_a, describes the competing possibilities. A null hypothesis need not assert “no effect”: it may specify a threshold or an inequality. (itl.nist.gov)

For example, a two-sided test of a population mean compares H0:μ=μ0H_0:\mu=\mu_0 with H1:μ≠μ0H_1:\mu\ne\mu_0. A one-sided test instead targets departures in a specified direction. A simple hypothesis completely specifies the data’s probability distribution, whereas a composite hypothesis permits more than one distribution. Hypotheses and assumptions jointly determine how a random variable representing the data behaves under the null model. (stats.org.uk)

Test statistics and decision rules

A test statistic reduces the observations to a quantity relevant to the hypotheses. Its sampling distribution describes its behavior across hypothetical repetitions of the sampling procedure. A rejection region contains outcomes that lead to rejection of H0H_0; its boundaries are often expressed as critical values. (itl.nist.gov)

The significance level α\alpha specifies an upper bound on the probability of rejecting a true null hypothesis, under the model. More formally, if RR is the rejection region and Θ0\Theta_0 the set of parameter values allowed by H0H_0, a level-α\alpha test satisfies

sup⁡θ∈Θ0Pθ(X∈R)≤α.\sup_{\theta\in\Theta_0}P_\theta(X\in R)\leq\alpha.

The bound need not be attained exactly, particularly for discrete data. The significance level is a property of the procedure, not the probability that a particular conclusion is wrong. (stats.org.uk)

A p-value is the probability, under the specified null model, of obtaining a test statistic at least as extreme as the observed value, with “extreme” defined by the test. A conventional decision rule rejects when p≤αp\leq\alpha. The p-value is neither the probability that H0H_0 is true nor the probability that the observations arose “by chance alone.” (doi.org)

Errors and statistical power

A Type I error occurs when a true null hypothesis is rejected. A Type II error occurs when a false null hypothesis is not rejected. For a particular alternative parameter value θ\theta, its probability is conventionally denoted β(θ)\beta(\theta). The power at that alternative is

π(θ)=1−β(θ)=Pθ(reject H0).\pi(\theta)=1-\beta(\theta) =P_\theta(\text{reject }H_0).

Power therefore varies across alternatives rather than being a single universal attribute of a test. (itl.nist.gov)

Power depends on sample size, variability, the magnitude of the departure from the null, and the rejection rule. Larger samples generally make smaller departures detectable. Experimental design and sample-size planning can incorporate power against specified alternatives. Failure to reject may reflect limited information, rather than agreement with the null hypothesis. (stat.berkeley.edu)

Example: testing a population mean

Suppose X1,…,XnX_1,\ldots,X_n are independent observations from a normal distribution with unknown mean and variance. The one-sample Student’s t-test of H0:μ=μ0H_0:\mu=\mu_0 uses

T=Xˉ−μ0S/n,T=\frac{\bar X-\mu_0}{S/\sqrt n},

where Xˉ\bar X is the sample mean and SS the sample standard deviation. The denominator estimates the standard error of the mean. Under H0H_0, TT follows Student’s t-distribution with n−1n-1 degrees of freedom. (itl.nist.gov)

For a two-sided level-α\alpha test, rejection occurs when

∣T∣>t1−α/2,n−1.|T|>t_{1-\alpha/2,n-1}.

With α=0.05\alpha=0.05, hypothetical values n=25n=25, Xˉ=102\bar X=102, S=5S=5, and μ0=100\mu_0=100 give T=2T=2. The critical value is approximately 2.0642.064, so the null is not rejected. This is not evidence that the population mean is exactly 100. (itl.nist.gov)

Two-sample tests compare population means. Welch’s formulation accommodates unequal population variances, while paired tests analyze within-pair differences rather than treating all observations as unrelated. Analysis of variance extends mean comparisons to several groups. (itl.nist.gov)

Confidence intervals and multiple comparisons

Tests and confidence intervals have a mathematical duality. Inverting a family of level-α\alpha tests produces a confidence set containing parameter values not rejected by those tests. For matching two-sided t procedures, a hypothesized mean is rejected exactly when it lies outside the corresponding 1−α1-\alpha confidence interval. (stat.berkeley.edu)

The multiple comparisons problem arises when many tests are considered together: individual error control does not automatically provide equivalent control for the collection. The Bonferroni correction addresses this by testing each of mm hypotheses at level α/m\alpha/m, ensuring that the probability of at least one Type I error is at most α\alpha, without requiring independence. (itl.nist.gov)

Interpretation and reporting

Statistical significance does not measure effect size or practical importance. A small effect can be statistically significant in a large sample, while an important effect may remain undetected in a small sample. Estimates and uncertainty intervals provide information that a rejection decision alone cannot convey. (itl.nist.gov)

Interpretation also depends on model assumptions, study execution, and the analyses actually performed. Selectively reporting only small p-values obscures the testing process and undermines their interpretation. The American Statistical Association’s 2016 statement emphasized transparent reporting and cautioned against basing scientific conclusions solely on whether a p-value crosses a threshold. (doi.org)