A/B TestingStatistics

Reading A/B Test Results Without Fooling Yourself

The statistical traps that make losing tests look like winners, and a simple discipline for calling results you can trust.

August 30, 2026 7 min read

Run It Right

Check Balance

Wait Long Enough

Call It Honestly

The hardest part of A/B testing is not running the test - it is not lying to yourself about the result. Peeking, tiny samples, and ignored guardrail metrics turn noise into false confidence, and false confidence ships changes that quietly cost money.

1"95% Significant" Is Not a Finish Line

Statistical significance tells you the result is unlikely to be pure chance *given the data so far*. It says nothing about whether you collected enough data, or whether you looked too early.

Peeking inflates false positives

If you check the dashboard daily and stop the moment it hits 95%, your real false-positive rate is closer to 25-30%, not 5%.

2Decide the Rules Before You Launch

  1. Pick the primary metric. One. Everything else is a guardrail or secondary.
  2. Estimate the minimum detectable effect you would care about.
  3. Use a sample-size calculator to get the required visitors per variant.
  4. Convert that to a duration in whole business weeks, and commit to it.

PRO TIP

Write these four numbers in the test ticket before it goes live:

  • Primary metric
  • Minimum effect worth shipping
  • Visitors per variant needed
  • End date (whole weeks)

3Watch the Guardrail Metrics

A variant can lift clicks and wreck revenue. Always read the whole picture.

If primary goes up but...It might mean
Revenue per visitor is flatYou shifted clicks, not sales
Refund rate risesThe copy over-promised
Add-to-cart up, checkout downNew friction later in the funnel
Only mobile improvedSegment before you roll out to everyone

4Check the Split Actually Happened

A 50/50 test that delivered 54/46 has a sample ratio mismatch. Something is broken - a redirect, a bot filter, an analytics gap - and the result is not trustworthy, however pretty the graph.

Quick gut check

With tens of thousands of visitors, a genuine 50/50 split stays within roughly a percentage point. Bigger gaps need investigation, not interpretation.

5Calling the Result

Ship it

Hit the pre-planned sample, significant on primary, guardrails clean.

Re-run it

Promising but under-powered, or a data issue mid-flight.

Kill it

Flat or negative at full sample. A clean null is a real, useful result.

Segment then decide

Strong on one device or audience, neutral elsewhere.

Half our "winners" stopped winning once we forced ourselves to wait for the planned sample size.Growth lead, subscription brand

FAQ

Only to stop a clearly harmful variant, or if the tool uses a sequential/Bayesian method designed for continuous monitoring. Not to lock in a win.
That is an answer: the effect is smaller than you hoped. Ship the cheaper variant or move on.
Industry-wide, roughly 1 in 5 to 1 in 8. If most of your tests win, your analysis is too loose.

Want a testing program you can trust?

We design, run, and analyse experiments for e-commerce teams.

Book a Free Consultation