Experiment metrics: primary, proxy, guardrail
The pizzeria launched a "second pizza −50%" promo: orders grew — and the margin? Half of A/B failures are not in the statistics but in WHAT was measured. One metric almost never describes an effect honestly: something grew while something else quietly sagged. So an experiment's design includes not one metric but a coordinated set with distinct roles.
The decision is made on ONE primary metric — but only if the guardrail metrics have not dipped. A proxy speeds up the test, but proxy growth without primary growth is a trap. Informational metrics explain the "why", but decisions are not made on them.
The primary metric (OEC) — one, directly reflecting the hypothesis and the goal. Changing the checkout button → the primary metric is "conversion to order". The decision is made on it, and the sample size is computed for it. If there are several primary metrics pulling in different directions, the decision becomes arbitrary.
Choosing the primary metric is itself the decision: what you are willing to degrade and what you are not. Without guardrails a "win" can silently be paid for with churn.
Good design always starts with "on which ONE metric will we decide, and which metrics will we not allow to sink". Without guardrails it is easy to ship a change that lifted clicks but cut revenue or burned retention.
Proxy metrics are a powerful accelerator and the main source of self-deception: the team happily moves the proxy while the real goal stands still. So the proxy→goal link is checked on historical data, not taken on faith.
Large products keep a set of standard guardrails (load speed, crashes, complaints, unsubscribes) checked automatically in EVERY experiment, whatever its topic.
Search and feed products are especially careful with proxies: "time in app" is easily inflated with addictive junk, so long-term retention always sits beside it as a guardrail.