What this calculator does
This compares two independent proportions. Groups of 100 with 45 and 30 successes give z = 2.191 and a p-value of 0.0285, significant at the 5% level.
Sample size does all the work here. The identical 45% against 40% comparison is not significant at n = 100 per group, with p = 0.4745, but is significant at n = 1,000 per group, with p = 0.0237. The difference between the groups never changed; only the precision with which it was measured did. That is worth remembering before treating significance as a measure of importance.
The formula
The two proportions are compared against a pooled estimate combining both samples, which is the correct standard error under the null hypothesis that they are equal. The z-statistic is the difference in proportions divided by that pooled standard error, and the p-value is two-tailed. The normal approximation requires a reasonable number of successes and failures in each group.
| Term | Meaning |
|---|---|
| Pooled proportion | The combined success rate across both groups, used for the standard error under the null. |
| Two-tailed | Testing for a difference in either direction. |
| Independence | The two samples must not overlap or be paired. For paired binary data, use McNemar test. |
| Statistical significance | A statement about precision, not about the size or importance of the difference. |
The inputs explained
| Field | What to enter |
|---|---|
| Group 1 successes | Successes in group 1. |
| Group 1 size | Size of group 1. |
| Group 2 successes | Successes in group 2. |
| Group 2 size | Size of group 2. |
When to use it
Comparing conversion rates
The standard A/B test comparison, where two variants are shown to independent groups.
Comparing treatment and control
Whether an event rate differs between two arms of a trial.
Comparing two survey groups
Whether two demographics answer a yes-or-no question differently.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How does sample size change the verdict?
The same comparison at different group 1 success counts.
| Group 1 successes (of 100) | Z-statistic | Two-tailed p-value | Proportion 1, 2 |
|---|---|---|---|
| 40 of 100 | 0 | 1.0000 | 40.0%, 40.0% |
| 45 of 100 | 0.7152 | 0.4745 | 45.0%, 40.0% |
| 50 of 100 | 1.421 | 0.1552 | 50.0%, 40.0% |
| 55 of 100 | 2.124 | 0.0337 | 55.0%, 40.0% |
Questions
Why is my difference not significant?
Most often because the sample is too small to detect a difference of that size, not because no difference exists. A five-point gap needs roughly 1,000 per group to be reliably detected. Not significant means not demonstrated, which is different from shown to be absent.
Why use a pooled proportion?
Because the test assumes the two proportions are equal under the null hypothesis, so the best estimate of that common value uses all the data from both groups. Using separate standard errors would be inconsistent with the hypothesis being tested.
Can I use this for a before-and-after comparison?
No, not if the same people are measured twice. That is paired data and requires McNemar test, which uses only the cases that changed. Treating paired data as independent throws away the pairing and gives the wrong answer.
Should I report the effect size as well?
Yes. A p-value says only whether a difference was detectable. The difference in proportions, and ideally a confidence interval around it, says how large it is. With a very large sample, a trivial difference will be significant, which is why both figures are needed.
For a single proportion against a target, see the z-test for a proportion calculator. For paired binary outcomes, see the McNemar’s test calculator.