What this calculator does
Running many tests at the 5% level makes a false positive nearly inevitable. Across 20 comparisons, the chance of at least one false positive is 64.2%, and across 50 it is 92.3%.
The Bonferroni correction divides the significance level by the number of comparisons, so 20 tests each use 0.25% rather than 5%. It is the simplest fix and the most conservative, guaranteeing the family-wise error rate stays below the stated level at the cost of making every individual test harder to pass.
The formula
The adjusted threshold is alpha divided by the number of comparisons. The chance of at least one false positive without correction is one minus the probability of avoiding one on every test, which is 1 − (1−α)^m. That figure climbs quickly and is the reason the correction exists.
| Term | Meaning |
|---|---|
| Family-wise error rate | The chance of at least one false positive across the whole set of tests. |
| Adjusted alpha | α divided by the number of comparisons. |
| Conservative | Erring toward not rejecting, which reduces false positives at the cost of power. |
| False discovery rate | An alternative framework controlling the proportion of false positives among rejections rather than the chance of any. |
The inputs explained
| Field | What to enter |
|---|---|
| Family-wise significance level (α) (%) | The family-wise significance level you want to maintain across all tests, usually 5%. |
| Number of comparisons (m) | The number of comparisons being made. Count every test performed, not just the ones reported. |
When to use it
Post hoc comparisons after ANOVA
Comparing every pair of groups produces many tests, and the correction keeps the overall error rate controlled.
Testing many outcomes
A trial measuring twenty endpoints will produce a significant one by chance alone unless corrected.
Subgroup analysis
Splitting results by many subgroups multiplies the tests, which is a well-known route to spurious findings.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How bad is the uncorrected error rate?
The adjusted threshold and uncorrected false positive risk for various test counts.
| Comparisons | Bonferroni-adjusted α per test | Chance of ≥1 false positive if uncorrected | Required p-value for significance |
|---|---|---|---|
| 1 test | 5.00% | 5.00% | 0.0500 |
| 5 tests | 1.00% | 22.6% | 0.0100 |
| 10 tests | 0.500% | 40.1% | 0.0050 |
| 20 tests | 0.250% | 64.2% | 0.0025 |
| 50 tests | 0.100% | 92.3% | 0.0010 |
Questions
Is Bonferroni too conservative?
Often, yes. It controls the family-wise error rate strictly, which means a real effect can easily be missed when there are many tests. The Holm-Bonferroni method is uniformly more powerful with the same guarantee, and there is rarely a good reason to prefer plain Bonferroni over it.
How many comparisons should I count?
Every test you performed, including those you decided not to report. Counting only the interesting results defeats the purpose entirely, since the inflated error rate comes from how many chances you gave yourself, not how many you chose to mention.
What is the alternative to Bonferroni?
Holm-Bonferroni for the same strict guarantee with more power, or Benjamini-Hochberg when you are willing to control the false discovery rate instead. The latter is standard in fields running thousands of tests, such as genomics, where family-wise control would leave nothing significant.
Does this apply to planned comparisons?
A small number of comparisons specified in advance is usually treated more leniently than exploratory testing, though opinions differ. What is not defensible is running many tests, finding one significant, and then claiming it was the plan all along.
For the test that precedes post hoc comparisons, see the one-way ANOVA calculator. For interpreting a single result, see the p-value calculator.