Contents

Data Analysis › Statistics

Distribution Shapes

Symmetry, tails and peaks — the shape that tells you which summary is honest.

Also known as: distribution shape, skew, data distribution

A distribution is just the picture of how often each value occurs. Its shape — symmetric or lopsided, one peak or two, heavy tail or none — decides which summary numbers are honest.

Two shapes cover most of what an analyst meets:

  • Symmetric, like the bell curve. The mean and the median land in the same place, and the standard deviation describes spread in a way that carries over into the normal distribution.
  • Skewed right, with a cluster of small values and a long tail of large ones. Order values, load times and salaries all look like this. The mean is dragged above the median, and one enormous value can move it a long way. This is the long tail.

Form the habit of looking before you summarise. Two columns can share a mean and a range and still be very different things — one of them is really two populations:

Column A: 40 41 42 43 44 45      mean 42.5, median 42.5
Column B: 12 15 30 55 70 73      mean 42.5, median 42.5

Column B is really two groups mixed together, which usually means you have two populations inside one column (segmentation) — and a pooled average can hide a trend that reverses when you split it (Simpson’s paradox).

Three readings worth doing every time:

  • One peak or two. Two peaks normally mean two populations were pooled. Split them before reporting a single average.
  • Where the tail is. A long right tail puts the mean above most of the data, so report the median and a high percentile next to it.
  • Where the gaps are. A sharp jump in the middle of the histogram is often a cap, a default, or a data error (data quality) rather than real behaviour.

A histogram is the fastest way to see shape; a box plot shows the same information with far less ink.