The Bonferroni correction is a method in statistics for controlling false positives when several hypothesis tests are considered together. For a family of tests and an overall significance level , it tests each hypothesis at level . Equivalently, each p-value is multiplied by , with the result capped at 1. The procedure controls the family-wise error rate: the probability of rejecting at least one true null hypothesis. Its validity does not require independence among the tests. (stat.ethz.ch)
Multiple testing and the decision rule
In multiple testing, a collection of individually valid tests can produce an appreciable probability of at least one Type I error. If all null hypotheses are true and the tests have independent rejection events, each with probability , the unadjusted family-wise error rate is
For example, 20 independent tests conducted at level 0.05 have an approximately 0.642 probability of at least one false rejection. This is a calculation under the stated assumptions, not a universal error rate for every collection of 20 tests. (stat.ethz.ch)
Bonferroni replaces the individual threshold with
For hypothesis , the decision rule is to reject when . The equivalent adjusted p-value is
which is compared with the original . These are alternative expressions of the same correction; applying both would impose an additional, unnecessary adjustment. (stat.ethz.ch)
As an illustrative calculation, suppose five tests yield p-values . At overall level 0.05, the threshold is 0.01, so only the first hypothesis is rejected. The adjusted values are , giving the same decision.
Mathematical basis
The correction follows from the union bound, also called Boole’s inequality. For events ,
This inequality places no requirement on statistical independence or on the events’ dependence structure. (webhome.auburn.edu)
Let denote the indices of the true null hypotheses. A valid null p-value satisfies for every . Applying the union bound to the false-rejection events gives
This derivation explains strong control: the bound holds regardless of which hypotheses are true, rather than only when every null hypothesis is true. It also shows why an invalid individual test is not repaired merely by adjusting its p-value. (arxiv.org)
Simultaneous confidence intervals
The same probability argument applies to confidence intervals. If each of intervals has coverage at least , the probability that all intervals simultaneously cover their respective parameters is at least . Thus, ten intervals intended to provide 95% simultaneous coverage can each be constructed with 99.5% individual coverage. (itl.nist.gov)
For a two-sided interval based on Student’s t-distribution, a common form is
where is an estimate, its standard error, and the relevant degrees of freedom. The factor 2 allocates the individual error probability between the two tails. Bonferroni intervals are therefore wider than corresponding unadjusted intervals. (itl.nist.gov)
Defining the family
The number refers to the hypotheses covered by the joint error-control claim, not automatically to the number of observations or groups. In a study comparing all pairs among group means, there are comparisons; comparing each noncontrol group with a single control involves only . The research question and experimental design determine which comparisons belong together. (itl.nist.gov)
Bonferroni is applicable to a preselected collection of contrasts in analysis of variance, rather than being restricted to all pairwise comparisons. A guarantee for one such collection does not automatically extend to additional comparisons outside it. This follows from the proof: the error bound covers only the rejection events included in the union. (itl.nist.gov)
Conservatism and alternatives
Bonferroni controls an upper bound, so the actual family-wise error rate may be below . Its conservatism can reduce statistical power—the ability to reject a false null in favor of an alternative hypothesis—especially when the family contains many tests. The computational simplicity and straightforward simultaneous intervals are nevertheless useful properties. (stat.ethz.ch)
The Holm–Bonferroni method orders the p-values and successively compares them with , stopping at the first nonrejection. It also provides strong family-wise error control under arbitrary dependence and rejects every hypothesis rejected by ordinary Bonferroni, sometimes more. (stat.ethz.ch)
Procedures controlling the false discovery rate, including the Benjamini–Hochberg procedure, address a different criterion: the expected proportion of false rejections among all rejections, with that proportion defined as zero when none occur. They do not generally provide the same probability bound on making any false rejection. (stat.ethz.ch)
A weighted extension assigns predetermined nonnegative weights summing to 1 and uses thresholds . The union-bound argument still gives family-wise error control for valid individual tests. Equal weights recover the ordinary correction. (pmc.ncbi.nlm.nih.gov)