Interactive statistics

56 free interactive lessons: from mean and median to A/B tests and Bayes. Drag the charts and build intuition. No sign-up.

Module 9: Data traps

Regression to the mean

The worst stores improved after a "performance review". The best employees did more modestly after a bonus. It looks like the intervention worked — but often it is an illusion, pure statistics. Its name is regression to the mean, and it fools even the experienced. Move the "share of luck" and watch the extremes roll back.

test 2 = test 1meanroll-back to the meantest 1 result (first time) →test 2 result (repeat) ↑top 20% by test 1
The best by test 1 (yellow zone): on average 1.69 on test 1, but already 0.85 on the repeat — they rolled back toward the mean (0) by themselves, though nothing was done to them.

Each point is the same object measured TWICE (say, a salesperson over two months or a student on two quizzes). Result = stable SKILL + random LUCK. The "share of luck" is how random the result is: 0% — pure skill (the repeat matches the first time, points on the gray diagonal); 100% — pure chance (the repeat is unrelated to the first). The red arrow shows the roll-back: the best by test 1 are on average lower on the repeat — because their luck comes out differently this time. Move the slider: the more luck, the stronger the roll-back.

Take any result that mixes skill and chance (sales, grades, KPIs). Each point is the same object measured TWICE: horizontally the first result (test 1), vertically the repeat (test 2). The gray diagonal means "the repeat exactly equals the first time": if the result were pure skill, every point would lie on it. The top 20% by test 1 are highlighted in yellow, and the red arrow shows where that group landed on the repeat.

What it means

This is one of the most common reasons for false faith in "working" measures. "The laggards took a training — they caught up", "we fined the worst branch — it shaped up": without a control group you are measuring regression to the mean, not the measure's effect.

The practical reflex: if a measure was applied to an extreme group and judged by "before/after" — be suspicious. You need a control or randomization (the evidence pyramid and A/B again).

Where it shows up

The "Sports Illustrated cover jinx" and the "rookie of the year curse": those who hit a peak then perform more modestly — not because of a jinx, but because luck helped throw them onto the peak.

In medicine, patients come in at their worst (the extreme) and then improve on their own — so without a control group any "treatment" looks effective. Hence the necessity of RCTs.

Definitions
Regression to the mean
extreme values are on average closer to the mean when re-measured — because of the random component of the result.
Skill + luck
result = a stable part (skill) + a random one (luck). It is the luck that rolls back, so the extremes "return".
Selection on the extreme
intervening on the very worst/best — where the regression-to-the-mean effect is largest and easiest to mistake for a result.
Next →
Enjoying the materials?

This site is built by the DataSlice Telegram channel: data analytics in plain words — real cases, metrics, careers. The channel is in Russian.

Open the Telegram channel