Skip to main content

How long do I have to run an ad test before I can trust the winner?

Google Ads Published August 7, 2026
Short Answer

Count conversions, not days or impressions. The smaller the difference you want to detect, the more conversions you need — and the requirement grows roughly fourfold each time you halve the effect size you care about. Most single-market service accounts will never accumulate enough conversions to resolve a small lift, which is an argument for testing big changes rather than headline word swaps.

The unit of a test is the conversion

Impressions and clicks feel like data because they accumulate quickly, but the thing you are trying to measure is the conversion rate, and the precision of that estimate depends on how many conversions you have observed. A test with many thousands of impressions and a handful of conversions per variant has told you almost nothing.

This is why so many ad tests get called early and reverse themselves later. The early leader in a low-conversion test is usually just the variant that got lucky first.

Why small lifts are effectively unmeasurable at local scale

Required sample scales with the inverse square of the effect size. Halving the difference you want to detect multiplies the conversions you need by about four. Detecting a large difference is cheap; detecting a modest one is expensive; detecting a small one is, at local service volumes, out of reach.

The honest conclusion is not that testing is pointless. It is that the tests worth running are the ones where you have a plausible reason to expect a large difference — a different offer, a different promise, a different call to action, a genuinely different page — rather than a comma.

Ways to buy yourself statistical power

  • Test at the highest level the change allows. An account-wide or campaign-group test aggregates conversions that a per-ad-group test would scatter.
  • Use a denser upstream event as a secondary read. Click-through rate resolves far faster than booking rate, and while it does not settle the question, a variant that loses badly on click-through rarely wins on bookings.
  • Run for whole cycles. Weekday and weekend demand differ enough in service businesses that a test spanning a partial week is measuring the calendar.
  • Do not restart the clock every time you tweak. Each edit resets the comparison, and accounts that are edited constantly can never finish a test.

What to do when the math says never

Many accounts simply cannot resolve the differences they care about. In that situation the right move is to stop pretending and use qualitative evidence deliberately: what callers say when they arrive, which objections they raise, which promises they repeat back. That is a smaller, faster, more honest signal than a statistically meaningless split test.

Reading it systematically across every call rather than the few a manager happens to hear is what call analysis contributes, and pairing it with what those calls were worth tends to produce better creative decisions than an underpowered experiment ever will.

Topics: ad testing · statistics · creative · low volume

Have a version of this question about your own business?

The useful answer usually depends on which systems you run and how they're connected. That's a conversation, not a blog post.

Related Answers

People who read this also asked

Browse the Answer Hub →

AI is easy to access. Making it useful is hard.

Bluefrog makes AI useful by integrating it with the way your business actually works — your software, your calls, your customers, your marketing and your revenue.

Technology development since 1997 · AI integration platforms since 2001