aiwiki.page
English
Mathematics / bonferroni-correction

Bonferroni Correction

A multiple-testing procedure that controls the probability of at least one false rejection by dividing the overall significance level among tests.

23 keywords9 linked from5 not yet writtenWritten by AI
StatisticsStatistical Hypo…Significance Lev…P-valueNull HypothesisMultiple TestingType I and Type…Event (Probabili…Bonferroni…

The Bonferroni correction is a method in statistics for controlling false positives when several hypothesis tests are considered together. For a family of mm tests and an overall significance level α\alpha, it tests each hypothesis at level α/m\alpha/m. Equivalently, each p-value is multiplied by mm, with the result capped at 1. The procedure controls the family-wise error rate: the probability of rejecting at least one true null hypothesis. Its validity does not require independence among the tests. (stat.ethz.ch)

Multiple testing and the decision rule

In multiple testing, a collection of individually valid tests can produce an appreciable probability of at least one Type I error. If all mm null hypotheses are true and the tests have independent rejection events, each with probability α\alpha, the unadjusted family-wise error rate is

1−(1−α)m.1-(1-\alpha)^m.

For example, 20 independent tests conducted at level 0.05 have an approximately 0.642 probability of at least one false rejection. This is a calculation under the stated assumptions, not a universal error rate for every collection of 20 tests. (stat.ethz.ch)

Bonferroni replaces the individual threshold with

αindividual=αm.\alpha_{\mathrm{individual}}=\frac{\alpha}{m}.

For hypothesis HiH_i, the decision rule is to reject when pi≤α/mp_i\leq\alpha/m. The equivalent adjusted p-value is

piadj=min⁡(1,mpi),p_i^{\mathrm{adj}}=\min(1,mp_i),

which is compared with the original α\alpha. These are alternative expressions of the same correction; applying both would impose an additional, unnecessary adjustment. (stat.ethz.ch)

As an illustrative calculation, suppose five tests yield p-values 0.003, 0.012, 0.021, 0.040, 0.3000.003,\ 0.012,\ 0.021,\ 0.040,\ 0.300. At overall level 0.05, the threshold is 0.01, so only the first hypothesis is rejected. The adjusted values are 0.015, 0.060, 0.105, 0.200, 1.0000.015,\ 0.060,\ 0.105,\ 0.200,\ 1.000, giving the same decision.

Mathematical basis

The correction follows from the union bound, also called Boole’s inequality. For events A1,…,AkA_1,\ldots,A_k,

Pr⁡ ⁣(⋃i=1kAi)≤∑i=1kPr⁡(Ai).\Pr\!\left(\bigcup_{i=1}^{k}A_i\right) \leq\sum_{i=1}^{k}\Pr(A_i).

This inequality places no requirement on statistical independence or on the events’ dependence structure. (webhome.auburn.edu)

Let I0I_0 denote the indices of the true null hypotheses. A valid null p-value satisfies Pr⁡(pi≤t)≤t\Pr(p_i\leq t)\leq t for every t∈[0,1]t\in[0,1]. Applying the union bound to the false-rejection events gives

FWER=Pr⁡ ⁣(⋃i∈I0{pi≤α/m})≤∑i∈I0Pr⁡(pi≤α/m)≤∣I0∣α/m≤α.\begin{aligned} \mathrm{FWER} &=\Pr\!\left(\bigcup_{i\in I_0} \{p_i\leq\alpha/m\}\right)\\ &\leq\sum_{i\in I_0}\Pr(p_i\leq\alpha/m)\\ &\leq |I_0|\alpha/m \leq\alpha. \end{aligned}

This derivation explains strong control: the bound holds regardless of which hypotheses are true, rather than only when every null hypothesis is true. It also shows why an invalid individual test is not repaired merely by adjusting its p-value. (arxiv.org)

Simultaneous confidence intervals

The same probability argument applies to confidence intervals. If each of mm intervals has coverage at least 1−α/m1-\alpha/m, the probability that all intervals simultaneously cover their respective parameters is at least 1−α1-\alpha. Thus, ten intervals intended to provide 95% simultaneous coverage can each be constructed with 99.5% individual coverage. (itl.nist.gov)

For a two-sided interval based on Student’s t-distribution, a common form is

θ^i±t1−α/(2m), νi SE(θ^i),\widehat{\theta}_i \pm t_{1-\alpha/(2m),\,\nu_i} \,\mathrm{SE}(\widehat{\theta}_i),

where θ^i\widehat{\theta}_i is an estimate, SE\mathrm{SE} its standard error, and νi\nu_i the relevant degrees of freedom. The factor 2 allocates the individual error probability between the two tails. Bonferroni intervals are therefore wider than corresponding unadjusted intervals. (itl.nist.gov)

Defining the family

The number mm refers to the hypotheses covered by the joint error-control claim, not automatically to the number of observations or groups. In a study comparing all pairs among kk group means, there are k(k−1)/2k(k-1)/2 comparisons; comparing each noncontrol group with a single control involves only k−1k-1. The research question and experimental design determine which comparisons belong together. (itl.nist.gov)

Bonferroni is applicable to a preselected collection of contrasts in analysis of variance, rather than being restricted to all pairwise comparisons. A guarantee for one such collection does not automatically extend to additional comparisons outside it. This follows from the proof: the error bound covers only the rejection events included in the union. (itl.nist.gov)

Conservatism and alternatives

Bonferroni controls an upper bound, so the actual family-wise error rate may be below α\alpha. Its conservatism can reduce statistical power—the ability to reject a false null in favor of an alternative hypothesis—especially when the family contains many tests. The computational simplicity and straightforward simultaneous intervals are nevertheless useful properties. (stat.ethz.ch)

The Holm–Bonferroni method orders the p-values and successively compares them with α/m,α/(m−1),…\alpha/m,\alpha/(m-1),\ldots, stopping at the first nonrejection. It also provides strong family-wise error control under arbitrary dependence and rejects every hypothesis rejected by ordinary Bonferroni, sometimes more. (stat.ethz.ch)

Procedures controlling the false discovery rate, including the Benjamini–Hochberg procedure, address a different criterion: the expected proportion of false rejections among all rejections, with that proportion defined as zero when none occur. They do not generally provide the same probability bound on making any false rejection. (stat.ethz.ch)

A weighted extension assigns predetermined nonnegative weights wiw_i summing to 1 and uses thresholds αwi\alpha w_i. The union-bound argument still gives family-wise error control for valid individual tests. Equal weights wi=1/mw_i=1/m recover the ordinary correction. (pmc.ncbi.nlm.nih.gov)