Contents

Data Analysis › Statistics

Non-Parametric Tests

Tests that fall back to ranks when the data does not meet the assumptions of a t-test.

Also known as: non-parametric test, distribution-free test, rank test

Non-parametric tests replace the values with their ranks and then run the test on the ranks. That one substitution buys a great deal of robustness and costs some sensitivity, and knowing which you are buying is the whole decision.

The ones an analyst reaches for most:

TestDesignQuestion it answers
Mann-Whitney Utwo independent groupsdo values in one group tend to be larger?
Wilcoxon signed-ranktwo paired groupsdo the paired differences tend one way?
Kruskal-Wallisthree or more independent groupsdoes at least one group tend to be larger?

Why you would use one: the metric is heavily skewed or carries extreme values (outliers); the sample is too small for a t test’s normal approximation to be credible; the data is ordinal — star ratings, survey scales — where the gap between adjacent values is not a real quantity; or the two distributions have clearly different shapes, so comparing means is misleading in the first place.

The trade-offs to state when you report one:

  • Sensitivity. On well-behaved, roughly normal data, the rank version typically needs more observations to detect the same difference. If your data passes the assumptions, the parametric test is the more sensitive choice.
  • What the null says. The Mann-Whitney U null is that a value from one group is no more likely to be larger than a value from the other, i.e. that the two distributions are similar. It is commonly described as a median test, which is not quite right: two groups can share a median and still differ enough for the test to detect. Report the medians and the IQR rather than claiming the test found a shift in the median.
  • Ties and zeros. Rank tests on data with many tied values — common with counts, where a large share of rows are zero — need the tie handling and the exact null distribution from your tool. Small samples should use exact p-values rather than the large-sample approximation, and which of the two you get is a software choice. Check it.
  • One-sided versus two-sided, and paired versus independent, all change the question. Consistency matters more than which you pick, and the decision rule should be fixed before the data.

A rank test is not a licence to skip the plot. Look at the distribution first: the reason the parametric assumptions failed is often the most interesting thing in the column.