Interactive statistics

56 free interactive lessons: from mean and median to A/B tests and Bayes. Drag the charts and build intuition. No sign-up.

Module 5: Hypothesis testing

Sample size and power

The previous lesson showed: more data, fewer errors. But how much exactly? You do not guess this — you compute it in advance. Otherwise you risk running an experiment physically incapable of noticing the effect.

critical value0MDE = +30 ₽ (3%)estimate of the between-group mean difference, ₽ →
H0: no effect H1: the effect is real■ α — false alarm (5%)■ β — missed effect (83%)■ power = 1−β (17%)
target 80%n = 20power ↑sample size n →
At the current n = 20 power is 17%. The 80% target needs n ≈ 275 per group (standardized effect d = 0.15).

The metric is the average receipt, 1000 ₽. We control not the mean difference directly (the business goal sets it) but the MDE — the minimum % lift we want to catch reliably; it sets the center of the blue H1 (+30 ₽). The top chart shows two distributions of the estimate: gray under "no effect" (H0), blue under "the effect is real" (H1). The threshold splits the axis into decisions: red right of the threshold under H0 is α (noise taken for effect), yellow left of it under H1 is β (a real effect overlooked). Smaller σ or bigger n narrow the bells → β falls, power grows. The bottom chart folds this into the power-vs-n curve: n is chosen to fit the MDE, not the other way around.

First, where power comes from at all. The top chart shows two distributions of our estimate of the between-group difference. The gray bell on the left — the world with NO effect (H0): the difference is zero on average, but the estimate wanders around zero due to chance. The blue bell on the right — the world with a real effect (H1): the estimate wanders around the true difference. Horizontal axis — what difference we might measure.

What it means
What decision this changes

Sample size is a decision made before launch. Starting an underpowered test means agreeing in advance to miss the effect and waste the traffic.

Before launching an A/B, product analysts compute the sample size: "to notice a 1-percentage-point conversion lift with 80% power we need this many users per variant". Without that calculation the test easily turns out meaningless.

And conversely: on a huge sample even a negligible, business-useless difference becomes statistically significant. So you check not only significance but the effect size.

Where it shows up

Clinical trials size their patient counts exactly this way — undersampling means a working drug may be rejected merely for lack of data.

Sociologists size polls for the required precision; engineers size the number of load-test runs. Everywhere the data budget is computed first and collected second.

Definitions
Power
the probability of detecting an effect if it really exists (1 minus the probability of a miss).
Effect size
how large a difference we want to notice. A small effect demands a large sample.
Sample planning
computing the required n in advance so the test has the strength to notice the effect that matters.
Sensitivityd = разница / σ
a test's ability to notice small effects. Grows with effect size and sample, falls with the spread σ.
When the method lies (assumptions)

A sample-size calculation is only as honest as the expected-effect estimate behind it. That estimate is usually optimistic; the real effect is smaller — and the "properly sized" test ends up underpowered.

The metric's variance is taken from history, and it drifts: season, traffic mix, promotions. Build in a margin and recompute on fresh data — last quarter's σ can betray you.

Next →
Enjoying the materials?

This site is built by the DataSlice Telegram channel: data analytics in plain words — real cases, metrics, careers. The channel is in Russian.

Open the Telegram channel