In statistics, the significance level, conventionally denoted by , is a prespecified upper bound on the probability that a statistical hypothesis test rejects its null hypothesis when that hypothesis is true. It controls the first kind of error distinguished in Type I and Type II errors. Common choices are 0.05, 0.01, and 0.10. The level belongs to the testing procedure: it is not the probability that a particular conclusion is incorrect. (itl.nist.gov)
Mathematical definition
Suppose observations have a probability distribution , with parameter . The null hypothesis specifies a set of possible parameter values ; the alternative hypothesis specifies competing values. A nonrandomized test rejects the null when falls in a rejection region . It is a level- test if
The test’s size is its largest null rejection probability:
Thus, size is an actual property of the procedure, whereas its stated level is a bound. For a null specifying a single distribution, the supremum reduces to one probability. For a composite null, the bound must hold throughout . A test whose size is below its nominal level is conservative. (stat210a.berkeley.edu)
For a randomized test, a function gives the probability of rejection after observing . Its rejection probability under is the expected value , and the level condition becomes . Randomization can permit exact error control when discrete observations prevent an ordinary rejection region from attaining the desired size. (stat210a.berkeley.edu)
Critical values and p-values
Tests commonly reduce observations to a test statistic, whose sampling distribution under the null determines critical values. These values mark the rejection region. In an upper-tailed test, unusually large statistics trigger rejection; in a lower-tailed test, unusually small statistics do so. A two-sided test considers departures in both directions. (online.stat.psu.edu)
A p-value expresses how extreme the observed statistic is relative to the null model. For a compatible family of tests, the decision rule is commonly written
The distinction is fundamental: is chosen before examining the result, whereas is calculated from the observations. For example, an observed meets a 0.05 threshold but not a 0.01 threshold. This comparison does not assign a probability of truth to either hypothesis. (itl.nist.gov)
For a statistic having a standard normal distribution under the null, a conventional two-sided test at rejects when exceeds approximately 1.96, allocating 0.025 to each tail. An upper-tailed test at the same level uses approximately 1.645. These illustrate how the alternative’s direction changes the critical value without changing the total nominal error bound. (online.stat.psu.edu)
Repeated-sampling interpretation
The error bound concerns repeated application of a specified procedure under a true null and valid model assumptions. A test of size 0.05 would reject approximately 5% of the time over many repetitions in that setting. It does not imply that 5% of all significant findings are false: that proportion also depends on which hypotheses are true and how effectively the tests detect alternatives. (itl.nist.gov)
Similarly, failure to reject means that the observations did not cross the chosen rejection threshold. It does not establish that the null is true. The test may have limited ability to detect a departure, especially when that departure is small. (itl.nist.gov)
Choice of level and statistical power
The significance level is part of experimental design. Choosing it involves the consequences of false rejection and missed detection rather than a universal mathematical preference for 0.05. With the same sample size and nested rejection regions, lowering reduces false rejections but also reduces statistical power, the probability of rejecting under a specified alternative. Power also depends on effect size, sample size, and variability. (itl.nist.gov)
Nominal error control is conditional on the procedure’s assumptions. Dependence between observations, inappropriate distributional assumptions, or analysis choices influenced by observed results can undermine it. Specifying the analysis and threshold beforehand, potentially through preregistration, distinguishes planned decisions from exploratory ones. Merely labeling a threshold “0.05” does not establish that the complete analysis has a 5% false-rejection rate. (stat.berkeley.edu)
Confidence intervals
There is a close relationship between significance levels and confidence intervals. A confidence set constructed by inverting level- tests contains precisely the parameter values those tests do not reject. Its coverage is at least , subject to the same assumptions. (stat210a.berkeley.edu)
For the usual two-sided test of a population mean, rejection at 0.05 corresponds to the hypothesized mean lying outside the matching 95% interval. When population variability is estimated, the calculation commonly uses Student’s t-distribution, appropriate degrees of freedom, the sample mean, and its standard error. The equivalence requires matching procedures; it is not guaranteed for arbitrary tests and intervals. (itl.nist.gov)
Multiple testing and interpretation
In multiple testing, a per-test significance level does not generally control the probability of any false rejection across the entire family. The Bonferroni correction tests each of hypotheses at , bounding that family-wide probability by , without requiring statistical independence. Procedures controlling the false discovery rate target a different quantity: the expected proportion of false rejections among all rejections. (itl.nist.gov)
Statistical significance does not measure practical importance, effect magnitude, or reproducibility. A large sample can make a small effect significant, while a small sample can leave an important effect undetected. The American Statistical Association’s 2016 statement distinguishes p-values from probabilities that hypotheses are true and emphasizes that scientific conclusions cannot rest solely on crossing a threshold. (online.stat.psu.edu)