Evidential strength: why A/B at all
"We updated the menu on the pizzeria's site and orders grew" — is that proof? Yes and no: it all depends on HOW the observation was obtained. Evidence has a hierarchy of strength. Before dissecting A/B mechanics, let's see why it sits at the top of this pyramid. Click through the levels.
Not all arguments are equal. At the pyramid's bottom are the weakest: an opinion, a single case, a testimonial ("feels better to me"). They bend easily to biases and survivorship. The higher the level, the more reliable the conclusion and the harder it is to explain away by chance or a third cause.
The level of evidence determines which decisions it can carry: observations are fine for hypotheses, big bets demand an experiment.
The practical skill is asking of any conclusion: "which level of the pyramid produced it?". "The manager is sure" — the bottom. "There is a correlation in the data" — a bit higher. "We did before/after" — higher still. "We ran an honest A/B" — the top. It instantly calibrates trust in the decision.
A/B is expensive: it needs traffic, time, infrastructure. So it is reserved for important and contested decisions, while obvious trifles ship without it. But when the cost of error is high — only a randomized experiment gives a reliable answer.
Medicine has the same pyramid: expert opinion at the bottom, randomized controlled trials and their meta-analyses at the top. A drug will not be approved on observations — an RCT, the direct analog of A/B, is required.
In products, loud failures are often born at the pyramid's bottom: "it's obviously better this way" — shipped without a test — the metric sank. A/B exists precisely to check such "obvious" things.