A/B TestingStatistics
Reading A/B Test Results Without Fooling Yourself
The statistical traps that make losing tests look like winners, and a simple discipline for calling results you can trust.
Run It Right
Check Balance
Wait Long Enough
Call It Honestly
The hardest part of A/B testing is not running the test - it is not lying to yourself about the result. Peeking, tiny samples, and ignored guardrail metrics turn noise into false confidence, and false confidence ships changes that quietly cost money.
1"95% Significant" Is Not a Finish Line
Statistical significance tells you the result is unlikely to be pure chance *given the data so far*. It says nothing about whether you collected enough data, or whether you looked too early.
Peeking inflates false positives
If you check the dashboard daily and stop the moment it hits 95%, your real false-positive rate is closer to 25-30%, not 5%.
2Decide the Rules Before You Launch
- Pick the primary metric. One. Everything else is a guardrail or secondary.
- Estimate the minimum detectable effect you would care about.
- Use a sample-size calculator to get the required visitors per variant.
- Convert that to a duration in whole business weeks, and commit to it.
PRO TIP
Write these four numbers in the test ticket before it goes live:
- Primary metric
- Minimum effect worth shipping
- Visitors per variant needed
- End date (whole weeks)
3Watch the Guardrail Metrics
A variant can lift clicks and wreck revenue. Always read the whole picture.
| If primary goes up but... | It might mean |
|---|---|
| Revenue per visitor is flat | You shifted clicks, not sales |
| Refund rate rises | The copy over-promised |
| Add-to-cart up, checkout down | New friction later in the funnel |
| Only mobile improved | Segment before you roll out to everyone |
4Check the Split Actually Happened
A 50/50 test that delivered 54/46 has a sample ratio mismatch. Something is broken - a redirect, a bot filter, an analytics gap - and the result is not trustworthy, however pretty the graph.
Quick gut check
With tens of thousands of visitors, a genuine 50/50 split stays within roughly a percentage point. Bigger gaps need investigation, not interpretation.
5Calling the Result
Ship it
Hit the pre-planned sample, significant on primary, guardrails clean.
Re-run it
Promising but under-powered, or a data issue mid-flight.
Kill it
Flat or negative at full sample. A clean null is a real, useful result.
Segment then decide
Strong on one device or audience, neutral elsewhere.
Half our "winners" stopped winning once we forced ourselves to wait for the planned sample size.— Growth lead, subscription brand
FAQ
Only to stop a clearly harmful variant, or if the tool uses a sequential/Bayesian method designed for continuous monitoring. Not to lock in a win.
That is an answer: the effect is smaller than you hoped. Ship the cheaper variant or move on.
Industry-wide, roughly 1 in 5 to 1 in 8. If most of your tests win, your analysis is too loose.
Want a testing program you can trust?
We design, run, and analyse experiments for e-commerce teams.
Book a Free Consultation





