Interactive statistics

56 free interactive lessons: from mean and median to A/B tests and Bayes. Drag the charts and build intuition. No sign-up.

Module 6: Experiments: A/B

The A/B process step by step: from hypothesis to rollout

The pizzeria's hypothesis: "a new banner on the homepage will lift orders". An A/B test is not "switch it on and watch the number". It is a five-stage process, and most failures happen not in the statistics but in skipped design and validation steps. Click the stages β€” each unfolds its key questions and typical mistakes.

Phase 1. Business framing
  • βœ“The problem and the product hypothesis are stated
  • βœ“The scale and direction of the primary-metric change are understood
  • βœ“The economics are computed (the break-even point)
  • βœ“Risks and constraints assessed; a success criterion and an action plan are formulated

This is an abridged map; a real checklist has dozens of items. What matters is the order: business and design first, launch only after.

Stage 1 β€” the business framing, before any statistics. What problem are we solving? What is the product hypothesis? By how much and where should the primary metric move, and does it pay off (the break-even point)? What are the risks and constraints? Without a clear success criterion and an action plan the test is pointless: nobody knows what counts as a win.

What it means
What decision this changes

The process pins the decisions BEFORE the data: the metric, the MDE, the duration. Anything decided after peeking at results is no longer a decision β€” it is a rationalization.

Mature teams treat A/B as a checklist: stages in order, and no launch until the design is closed. It is not bureaucracy β€” every skipped item turns into either an invalid result or an unshipped conclusion.

Special attention to stage 2: metrics, sample size and the statistical test are fixed in writing BEFORE launch. That guards against bending the analysis toward the desired result (p-hacking) and makes the experiment reproducible.

Where it shows up

Big product companies run an "experimentation culture" with a platform and design reviews: peers check the hypothesis, metrics and sample calculation before the start. It sharply cuts the share of useless and faulty tests.

Clinical trials are built just as strictly: the protocol (the design's analog) is registered in advance, metrics and sample size are fixed, stopping follows the rules. The cost of error is high β€” so process beats intuition.

Definitions
Guardrail metrics
protective metrics that must not sink for the primary's sake (speed, errors, unsubscribes). Fixed at design time.
MDE
the minimum effect the test must be able to notice. It sets the required sample size.
SRM (Sample Ratio Mismatch)
the actual group ratio diverging from the expected one β€” a frequent sign of a split or logging bug.
A/A check
running the methodology on two identical groups: "significance" there means the test statistic is broken.
Next β†’
Enjoying the materials?

This site is built by the DataSlice Telegram channel: data analytics in plain words β€” real cases, metrics, careers. The channel is in Russian.

Open the Telegram channel