StatGardenREF. DESK
Calculators/Statistics/Mann-Whitney U test (Wilcoxon rank-sum)
Statistics

Mann-Whitney U test (Wilcoxon rank-sum) calculator

Compares two independent samples for a shift in distribution, using ranks instead of means.

Published 4 August 2026 · Updated 25 September 2026

What this calculator does

The Mann-Whitney U test, also called the Wilcoxon rank-sum test, compares two independent samples using only the ranks of the values. It makes no assumption that the data is normal, which is why it is the usual fallback when a t-test is not defensible.

The statistic counts how often a value from one sample exceeds a value from the other. Complete separation gives U = 0, which is the strongest possible result: with samples of 5 and 6, that produces z = −2.739 and p = 0.0062. Overlapping samples push U toward its midpoint of 15.

The formula

FormulaU₁ = R₁ − n₁(n₁+1)/2; μ_U = n₁n₂/2; σ_U² = n₁n₂/12·[(N+1) − Σ(t³−t)/(N(N−1))]; z = (U − μ_U)/σ_U

All values from both samples are pooled and ranked, with ties receiving average ranks. U is computed from the rank sum of each sample, and the smaller of the two U values is reported. For samples above about 8 each, the normal approximation with a tie correction gives the p-value; below that an exact table is more reliable.

TermMeaning
U statisticThe number of times a value in one sample exceeds one in the other.
Rank sumThe sum of ranks for a sample in the pooled ordering, from which U is derived.
NonparametricMaking no assumption about the shape of the distribution.
Tie correctionAn adjustment to the standard deviation of U when values are tied across samples.

The inputs explained

FieldWhat to enter
Sample ASample A values, comma or space separated.
Sample BSample B values. The two samples need not be the same size.

When to use it

Comparing small non-normal samples

Where a t-test assumption of normality cannot be justified and the sample is too small to rely on the central limit theorem.

Working with ordinal data

Ratings and ordered categories have ranks but no meaningful arithmetic, which suits a rank-based test.

Data with outliers

An extreme value becomes simply the highest rank, so it cannot dominate the test as it would a comparison of means.

Worked examples

Every figure in the tables below is produced by this page’s own calculator at build time, so the numbers and the tool always agree. Select any row to load that scenario.

How does separation affect the result?

Sample A ranging from overlapping to completely separated.

Sample B fixed at 9, 11, 14, 17, 20, 23
Sample AU statisticz (normal approximation)Two-tailed p-value
12, 15, 18, 22, 2510.00-0.91290.3613
20, 25, 30, 35, 401.50-2.4700.0135
1, 2, 3, 4, 50.00-2.7390.0062
The first row interleaves with sample B and gives U = 10 against an expected 15, nowhere near significant at p = 0.3613. The last row sits entirely below sample B, giving U = 0 and p = 0.0062. That is the smallest p-value achievable with these sample sizes, since no arrangement separates them further.

Questions

What does the Mann-Whitney test actually test?

Strictly, whether one sample tends to produce larger values than the other. It is often described as comparing medians, but that is only correct under the additional assumption that the two distributions have the same shape. Without that, it tests stochastic dominance rather than any specific parameter.

Is it the same as the Wilcoxon rank-sum test?

Yes, they are mathematically equivalent and differ only in how the statistic is expressed. The Wilcoxon signed-rank test is a different test entirely, used for paired data rather than two independent samples.

When should I use this instead of a t-test?

When the data is ordinal, when the sample is small and clearly non-normal, or when outliers would distort a mean comparison. For reasonably large samples from roughly normal data, a t-test has more power and remains preferable.

How does it handle ties?

Tied values receive the average of the ranks they span, and the standard deviation of U is adjusted downward to account for them. Heavy tying reduces the effective information in the data, which the correction reflects.

For comparing means directly, see the t-statistic calculator. For a rank-based association measure, see the Spearman’s rank correlation calculator.