Statistical hypothesis testing is a framework in statistics for evaluating claims about populations or data-generating processes using observed data. A test specifies competing hypotheses, a measure of discrepancy, and a rule for deciding whether the observations warrant rejecting a designated hypothesis. Its conclusions depend on the statistical model and sampling assumptions; they do not establish truth with certainty. (itl.nist.gov)
Hypotheses and statistical models
A statistical hypothesis specifies a property of a population or model, such as a mean, a proportion, or the relationship between variables. The null hypothesis, denoted , is the claim tested. The alternative hypothesis, denoted or , describes the competing possibilities. A null hypothesis need not assert “no effect”: it may specify a threshold or an inequality. (itl.nist.gov)
For example, a two-sided test of a population mean compares with . A one-sided test instead targets departures in a specified direction. A simple hypothesis completely specifies the data’s probability distribution, whereas a composite hypothesis permits more than one distribution. Hypotheses and assumptions jointly determine how a random variable representing the data behaves under the null model. (stats.org.uk)
Test statistics and decision rules
A test statistic reduces the observations to a quantity relevant to the hypotheses. Its sampling distribution describes its behavior across hypothetical repetitions of the sampling procedure. A rejection region contains outcomes that lead to rejection of ; its boundaries are often expressed as critical values. (itl.nist.gov)
The significance level specifies an upper bound on the probability of rejecting a true null hypothesis, under the model. More formally, if is the rejection region and the set of parameter values allowed by , a level- test satisfies
The bound need not be attained exactly, particularly for discrete data. The significance level is a property of the procedure, not the probability that a particular conclusion is wrong. (stats.org.uk)
A p-value is the probability, under the specified null model, of obtaining a test statistic at least as extreme as the observed value, with “extreme” defined by the test. A conventional decision rule rejects when . The p-value is neither the probability that is true nor the probability that the observations arose “by chance alone.” (doi.org)
Errors and statistical power
A Type I error occurs when a true null hypothesis is rejected. A Type II error occurs when a false null hypothesis is not rejected. For a particular alternative parameter value , its probability is conventionally denoted . The power at that alternative is
Power therefore varies across alternatives rather than being a single universal attribute of a test. (itl.nist.gov)
Power depends on sample size, variability, the magnitude of the departure from the null, and the rejection rule. Larger samples generally make smaller departures detectable. Experimental design and sample-size planning can incorporate power against specified alternatives. Failure to reject may reflect limited information, rather than agreement with the null hypothesis. (stat.berkeley.edu)
Example: testing a population mean
Suppose are independent observations from a normal distribution with unknown mean and variance. The one-sample Student’s t-test of uses
where is the sample mean and the sample standard deviation. The denominator estimates the standard error of the mean. Under , follows Student’s t-distribution with degrees of freedom. (itl.nist.gov)
For a two-sided level- test, rejection occurs when
With , hypothetical values , , , and give . The critical value is approximately , so the null is not rejected. This is not evidence that the population mean is exactly 100. (itl.nist.gov)
Two-sample tests compare population means. Welch’s formulation accommodates unequal population variances, while paired tests analyze within-pair differences rather than treating all observations as unrelated. Analysis of variance extends mean comparisons to several groups. (itl.nist.gov)
Confidence intervals and multiple comparisons
Tests and confidence intervals have a mathematical duality. Inverting a family of level- tests produces a confidence set containing parameter values not rejected by those tests. For matching two-sided t procedures, a hypothesized mean is rejected exactly when it lies outside the corresponding confidence interval. (stat.berkeley.edu)
The multiple comparisons problem arises when many tests are considered together: individual error control does not automatically provide equivalent control for the collection. The Bonferroni correction addresses this by testing each of hypotheses at level , ensuring that the probability of at least one Type I error is at most , without requiring independence. (itl.nist.gov)
Interpretation and reporting
Statistical significance does not measure effect size or practical importance. A small effect can be statistically significant in a large sample, while an important effect may remain undetected in a small sample. Estimates and uncertainty intervals provide information that a rejection decision alone cannot convey. (itl.nist.gov)
Interpretation also depends on model assumptions, study execution, and the analyses actually performed. Selectively reporting only small p-values obscures the testing process and undermines their interpretation. The American Statistical Association’s 2016 statement emphasized transparent reporting and cautioned against basing scientific conclusions solely on whether a p-value crosses a threshold. (doi.org)