Interactive statistics

56 free interactive lessons: from mean and median to A/B tests and Bayes. Drag the charts and build intuition. No sign-up.

Module 6: Experiments: A/B

Evidential strength: why A/B at all

"We updated the menu on the pizzeria's site and orders grew" — is that proof? Yes and no: it all depends on HOW the observation was obtained. Evidence has a hierarchy of strength. Before dissecting A/B mechanics, let's see why it sits at the top of this pyramid. Click through the levels.

↑ strongerweaker ↓
A/B test (randomized)
Random assignment to groups equalizes all other factors → the visible difference is caused by the change itself. The gold standard of causality.

Not all arguments are equal. At the pyramid's bottom are the weakest: an opinion, a single case, a testimonial ("feels better to me"). They bend easily to biases and survivorship. The higher the level, the more reliable the conclusion and the harder it is to explain away by chance or a third cause.

What it means
What decision this changes

The level of evidence determines which decisions it can carry: observations are fine for hypotheses, big bets demand an experiment.

The practical skill is asking of any conclusion: "which level of the pyramid produced it?". "The manager is sure" — the bottom. "There is a correlation in the data" — a bit higher. "We did before/after" — higher still. "We ran an honest A/B" — the top. It instantly calibrates trust in the decision.

A/B is expensive: it needs traffic, time, infrastructure. So it is reserved for important and contested decisions, while obvious trifles ship without it. But when the cost of error is high — only a randomized experiment gives a reliable answer.

Where it shows up

Medicine has the same pyramid: expert opinion at the bottom, randomized controlled trials and their meta-analyses at the top. A drug will not be approved on observations — an RCT, the direct analog of A/B, is required.

In products, loud failures are often born at the pyramid's bottom: "it's obviously better this way" — shipped without a test — the metric sank. A/B exists precisely to check such "obvious" things.

Definitions
Evidential strength
how reliably a piece of evidence establishes a causal link rather than a coincidence.
Randomization
random assignment to groups; it equalizes all other factors and thereby licenses causal conclusions.
Quasi-experiment
a comparison without randomization (before/after, different regions). Weaker than A/B: the effect can be confused with extraneous changes.
Next →
Enjoying the materials?

This site is built by the DataSlice Telegram channel: data analytics in plain words — real cases, metrics, careers. The channel is in Russian.

Open the Telegram channel