werbero.comThe magazine for advertising that works

Strategy & Craft

A/B tests that actually mean something

Most published test results would not survive a second run. Four conditions separate a genuine finding from a coincidence.

Mar 14, 2026 2 min read 407 words
A/B tests that actually mean something

Key points

  • Below a few hundred conversions per variant, most results are noise.
  • Stopping a test when it looks good is the most common way to reach a false conclusion.
  • Test one variable at a time or you learn nothing you can reuse.

A/B testing sounds objective and usually is not, because the arithmetic is skipped. A variant showing 18 per cent improvement on 40 conversions is not a finding; it is a coin landing heads twice.

Condition 1: sample size, calculated first

Before starting, calculate how many conversions per variant you need to detect the difference you care about. The smaller the effect, the larger the sample.

Baseline rateEffect to detectConversions needed per variant
2 %+20 % relativeroughly 3,900
5 %+20 % relativeroughly 1,500
10 %+20 % relativeroughly 700
10 %+50 % relativeroughly 130

Those numbers are the reason most small-site tests cannot work. A page with 800 visits a month and a 3 per cent conversion rate would need years.

Condition 2: do not stop early

The single most common error. A test is watched daily, and when one variant leads, it is declared the winner.

That procedure produces a false positive most of the time, because early results fluctuate wildly. Set the duration and sample size in advance and do not look at the result until both are reached.

If you would have stopped the test at the moment it looked good, the result tells you when you looked, not which variant is better.

Condition 3: one variable

Change the headline and the button and the image, and a winning variant tells you nothing transferable. You know this page is better; you do not know why, so you cannot apply it anywhere else.

Test one element. It is slower and it accumulates knowledge rather than isolated wins.

Condition 4: full weeks

Behaviour differs by day. A test running Tuesday to Friday measures a working-week audience. Always run in complete seven-day blocks, minimum two.

What is worth testing

Order by expected effect, not by ease.

  1. The offer itself. Free survey against fixed price against discount. Largest effects by far.
  2. The headline. Second largest.
  3. The form, specifically the number of fields.
  4. Proof elements, reviews and references above or below the fold.
  5. The call to action wording.
  6. Colours and small design details. Smallest effects, most frequently tested.

When not to test

With low traffic, testing is theatre. Two better alternatives exist.

Make changes based on evidence from elsewhere, since the large effects, such as page speed, form length and message match, are well established and do not need local proof.

Or test at business level: change one thing, watch total enquiries for three months against the previous year, and accept the lower precision in exchange for a measurable question.

Frequently asked questions

How long should a test run?

At least two full weeks, to cover both weekend and weekday behaviour, and until the sample size is reached.

What if traffic is too low to test?

Then do not run A/B tests. Make the changes you have good reasons for and measure at the business level.

More from Strategy & Craft

Keep reading

All articles
Splitting an ad budget: the 70-20-10 rule

Strategy & Craft

Splitting an ad budget: the 70-20-10 rule

Seventy per cent on what works, twenty on scaling, ten on experiments. The rule is old, simple, and violated in almost every account.

2 min readMay 9, 2026
The brief agencies actually need

Strategy & Craft

The brief agencies actually need

A bad brief produces bad work at full price, and the client pays twice. Eight questions, two pages, three hours of your time.

2 min readMay 5, 2026