What this calculator does
The Mann-Whitney U test, also called the Wilcoxon rank-sum test, compares two independent samples using only the ranks of the values. It makes no assumption that the data is normal, which is why it is the usual fallback when a t-test is not defensible.
The statistic counts how often a value from one sample exceeds a value from the other. Complete separation gives U = 0, which is the strongest possible result: with samples of 5 and 6, that produces z = −2.739 and p = 0.0062. Overlapping samples push U toward its midpoint of 15.
The formula
All values from both samples are pooled and ranked, with ties receiving average ranks. U is computed from the rank sum of each sample, and the smaller of the two U values is reported. For samples above about 8 each, the normal approximation with a tie correction gives the p-value; below that an exact table is more reliable.
| Term | Meaning |
|---|---|
| U statistic | The number of times a value in one sample exceeds one in the other. |
| Rank sum | The sum of ranks for a sample in the pooled ordering, from which U is derived. |
| Nonparametric | Making no assumption about the shape of the distribution. |
| Tie correction | An adjustment to the standard deviation of U when values are tied across samples. |
The inputs explained
| Field | What to enter |
|---|---|
| Sample A | Sample A values, comma or space separated. |
| Sample B | Sample B values. The two samples need not be the same size. |
When to use it
Comparing small non-normal samples
Where a t-test assumption of normality cannot be justified and the sample is too small to rely on the central limit theorem.
Working with ordinal data
Ratings and ordered categories have ranks but no meaningful arithmetic, which suits a rank-based test.
Data with outliers
An extreme value becomes simply the highest rank, so it cannot dominate the test as it would a comparison of means.
Worked examples
Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.
How does separation affect the result?
Sample A ranging from overlapping to completely separated.
| Sample A | U statistic | z (normal approximation) | Two-tailed p-value |
|---|---|---|---|
| 12, 15, 18, 22, 25 | 10.00 | -0.9129 | 0.3613 |
| 20, 25, 30, 35, 40 | 1.50 | -2.470 | 0.0135 |
| 1, 2, 3, 4, 5 | 0.00 | -2.739 | 0.0062 |
Questions
What does the Mann-Whitney test actually test?
Strictly, whether one sample tends to produce larger values than the other. It is often described as comparing medians, but that is only correct under the additional assumption that the two distributions have the same shape. Without that, it tests stochastic dominance rather than any specific parameter.
Is it the same as the Wilcoxon rank-sum test?
Yes, they are mathematically equivalent and differ only in how the statistic is expressed. The Wilcoxon signed-rank test is a different test entirely, used for paired data rather than two independent samples.
When should I use this instead of a t-test?
When the data is ordinal, when the sample is small and clearly non-normal, or when outliers would distort a mean comparison. For reasonably large samples from roughly normal data, a t-test has more power and remains preferable.
How does it handle ties?
Tied values receive the average of the ranks they span, and the standard deviation of U is adjusted downward to account for them. Heavy tying reduces the effective information in the data, which the correction reflects.
For comparing means directly, see the t-statistic calculator. For a rank-based association measure, see the Spearman’s rank correlation calculator.