StatGardenREF. DESK
Calculators/Statistics/Bonferroni correction
Statistics

Bonferroni correction calculator

Adjusts the significance threshold when running many statistical tests at once.

Published 8 August 2026 · Updated 25 September 2026

What this calculator does

Running many tests at the 5% level makes a false positive nearly inevitable. Across 20 comparisons, the chance of at least one false positive is 64.2%, and across 50 it is 92.3%.

The Bonferroni correction divides the significance level by the number of comparisons, so 20 tests each use 0.25% rather than 5%. It is the simplest fix and the most conservative, guaranteeing the family-wise error rate stays below the stated level at the cost of making every individual test harder to pass.

The formula

FormulaAdjusted α = α / m; probability of at least one false positive without correction = 1 − (1−α)ᵐ

The adjusted threshold is alpha divided by the number of comparisons. The chance of at least one false positive without correction is one minus the probability of avoiding one on every test, which is 1 − (1−α)^m. That figure climbs quickly and is the reason the correction exists.

TermMeaning
Family-wise error rateThe chance of at least one false positive across the whole set of tests.
Adjusted alphaα divided by the number of comparisons.
ConservativeErring toward not rejecting, which reduces false positives at the cost of power.
False discovery rateAn alternative framework controlling the proportion of false positives among rejections rather than the chance of any.

The inputs explained

FieldWhat to enter
Family-wise significance level (α) (%)The family-wise significance level you want to maintain across all tests, usually 5%.
Number of comparisons (m)The number of comparisons being made. Count every test performed, not just the ones reported.

When to use it

Post hoc comparisons after ANOVA

Comparing every pair of groups produces many tests, and the correction keeps the overall error rate controlled.

Testing many outcomes

A trial measuring twenty endpoints will produce a significant one by chance alone unless corrected.

Subgroup analysis

Splitting results by many subgroups multiplies the tests, which is a well-known route to spurious findings.

Worked examples

Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.

How bad is the uncorrected error rate?

The adjusted threshold and uncorrected false positive risk for various test counts.

α = 5%
ComparisonsBonferroni-adjusted α per testChance of ≥1 false positive if uncorrectedRequired p-value for significance
1 test5.00%5.00%0.0500
5 tests1.00%22.6%0.0100
10 tests0.500%40.1%0.0050
20 tests0.250%64.2%0.0025
50 tests0.100%92.3%0.0010
With a single test the correction does nothing, as it should. By 10 tests the uncorrected risk of at least one false positive is already 40.1%, and by 50 it is 92.3%. Running fifty uncorrected comparisons and reporting the significant ones is close to guaranteed to produce a finding, whether or not anything is there.

Questions

Is Bonferroni too conservative?

Often, yes. It controls the family-wise error rate strictly, which means a real effect can easily be missed when there are many tests. The Holm-Bonferroni method is uniformly more powerful with the same guarantee, and there is rarely a good reason to prefer plain Bonferroni over it.

How many comparisons should I count?

Every test you performed, including those you decided not to report. Counting only the interesting results defeats the purpose entirely, since the inflated error rate comes from how many chances you gave yourself, not how many you chose to mention.

What is the alternative to Bonferroni?

Holm-Bonferroni for the same strict guarantee with more power, or Benjamini-Hochberg when you are willing to control the false discovery rate instead. The latter is standard in fields running thousands of tests, such as genomics, where family-wise control would leave nothing significant.

Does this apply to planned comparisons?

A small number of comparisons specified in advance is usually treated more leniently than exploratory testing, though opinions differ. What is not defensible is running many tests, finding one significant, and then claiming it was the plan all along.

For the test that precedes post hoc comparisons, see the one-way ANOVA calculator. For interpreting a single result, see the p-value calculator.