ZenWeb - Blog - Google Ads Experiments: How to A/B Test Campaigns Right

Google Ads Experiments: How to A/B Test Campaigns Right

Jian Tat Lee
August 22, 2026

Share this post:

Google Ads Experiments: How to A/B Test Campaigns Right
TL;DR: Google Ads experiments split a campaign’s traffic between your current setup and a changed version, so you can see which one actually performs better. The setup takes ten minutes. The hard part is having enough conversions to read the result — below roughly 30 conversions a month, most experiments never produce a verdict you can trust. Test big changes, run them for at least four weeks, and ignore the winner Google shows you in week one.

Every Google Ads account has an argument in it somewhere. Should we move to Smart Bidding? Is the new landing page better? Does broad match help or bleed?

Most of these arguments get settled the same lazy way: change the thing, watch the numbers for a fortnight, and credit whatever happened next. That is not a test. That is a guess with a calendar attached.

Google Ads experiments exist to settle those arguments properly — by running both versions at the same time, against the same market, on the same days. This guide from ZenWeb covers what to test, how long to wait, how to read the result honestly, and — the part almost nobody says out loud — when your account is simply too small to test at all. New here? Start with how Google Ads works. Otherwise, here is the setup in a few minutes.

Running campaign experiments in Google Ads

Source video: How to Run Google Ads Campaign Experiments (10 Ideas + Fast Setup) on YouTube

1. What Is a Google Ads Experiment?

Quick Answer: A Google Ads experiment clones your existing campaign, applies the one change you want to test, and splits traffic between the two. Both versions run at the same time on the same budget pool, so seasonality, competitors and market swings hit both arms equally — and the difference you see is down to the change itself.

The critical word is simultaneous. A before-and-after comparison cannot separate your change from everything else that moved — a competitor pausing, a festive week, a price rise. A split test can, because both arms share the same conditions.

Google Ads experiments come in a few flavours, and the one you pick decides what you are allowed to change:

  • Custom experiments. The workhorse. Clone a Search or Display campaign, change almost anything — bidding, keywords, match types, audiences, landing pages — and split the traffic. Google’s custom experiment documentation walks through the setup screen.
  • Ad variations. A lighter tool for testing copy across many ads at once, without cloning the whole campaign.
  • Demand Gen A/B tests. Google’s Demand Gen experiment type tests creative and audience arms against each other — worth knowing if you already run Demand Gen campaigns alongside Search.
  • Performance Max uplift tests. Measure what PMax adds on top of your existing campaigns, rather than comparing two versions of the same thing.

For most Malaysian SME accounts, running Google Ads experiments means custom experiments on a Search campaign. That is the tool this guide focuses on.

Key takeaway: An experiment is not “change it and watch”. It is two versions running side by side, in the same week, for the same searches.

Not sure whether your account is testable yet?

Volume decides that, not ambition. See how ZenWeb manages Google Ads →


2. Which Changes Are Actually Worth Testing?

Quick Answer: Test big, structural changes — bidding strategy and landing page. Across ZenWeb-managed accounts these two produce a clear verdict about two-thirds of the time. Small changes, especially ad copy tweaks, rarely move enough traffic to be readable on an SME budget, and eat weeks of testing capacity to say nothing.

Every experiment costs something real: half your traffic spends the test period on a version you may end up throwing away. That is the tax. Only changes big enough to repay it are worth running.

Which Google Ads Experiments Produce a Usable Answer, by Test Type
Number of experiments run, share that reached a clear verdict, and median lift when the variant won, broken down by the type of change tested, across ZenWeb-managed Malaysian SME Google Ads campaigns from 2024 to 2026.
What was testedReached a clear verdictMedian lift when it won
Bidding strategy switch71%+18% conversions
Landing page swap66%+24% conversion rate
Keyword match type change52%+9% conversions
Audience or targeting change44%+7% conversions
Ad copy variation38%+6% clickthrough
Bid adjustment tweak34%+5% conversions

Source: Aggregated from ZenWeb-managed Google Ads campaigns, Malaysia, 2024–2026. Search campaigns only; “clear verdict” means one arm led consistently through the final two weeks of the test.

The pattern is blunt. The top two rows are worth your testing capacity; the bottom two mostly are not. That does not mean ad copy and bid adjustments do not matter — a split test is just the wrong instrument for them on a small budget. Improve those by judgement; save the experiment slot for arguments that deserve one.

The tests that repay the tax most reliably in Malaysian accounts:

Key takeaway: Test the arguments, not the details. Bidding and landing pages earn a verdict; headline tweaks usually just eat a month.

3. How Long Should an Experiment Run?

Quick Answer: Four weeks minimum, and longer if your base campaign converts fewer than 30 times a month. Duration is not a preference — it is set by your conversion volume. Halve the traffic, and a campaign getting 20 conversions a month is asking each arm to prove itself on ten.

This is where most Google Ads experiments quietly fail — not in the setup, but in the arithmetic. Split a low-volume campaign in two and each arm gets so few conversions that the difference is just luck.

Days to a Readable Result, by Base Campaign Conversion Volume
Typical number of days before a 50-50 Google Ads custom experiment produced a stable, readable verdict, grouped by the base campaign’s monthly conversion volume, across ZenWeb-managed Malaysian SME campaigns 2024 to 2026.
Conversions per month (base campaign)Days to a readable result
Under 15

Rarely readable

15–29

62 days

30–59

38 days

60–99

26 days

100 or more

17 days

Source: Aggregated from ZenWeb-managed Google Ads campaigns, Malaysia, 2024–2026. 50-50 custom experiments; “readable” means one arm led consistently for the final fourteen days.

Read the top row honestly. Under fifteen conversions a month, a split test usually hands you a number that looks like an answer and is not one. That is not a reason to give up on the account — it is a reason to improve it by reasoning rather than testing, and to spend the budget lifting conversion volume first.

Two timing rules save wasted months:

  1. Give Smart Bidding its learning period. A bidding experiment needs one to two weeks before the variant even behaves like itself. Judging it in week one judges the learning phase, not the strategy.
  2. Do not straddle a festive peak. Raya, Chinese New Year and year-end promotions distort both arms unevenly. If the test must run through one, set your seasonality adjustments first.
Key takeaway: Your conversion volume sets the test duration, not your patience. Under 15 conversions a month, do not split the campaign — grow it.

4. How to Set Up a Custom Experiment

Quick Answer: Open Campaigns → Experiments, create a custom experiment, pick your base campaign, change one thing in the copy, set a 50-50 split, and schedule at least four weeks. The one discipline that matters: change one variable. Two changes at once buys you a result you cannot act on.

The mechanics are quick. The prep matters.

  1. Check your conversion tracking first. Google Ads experiments are measured in conversions. If the tracking is wrong, the test is not useless — it is confidently wrong. Verify it against the conversion tracking setup.
  2. Pick the base campaign. One with real volume and a clean account structure. A messy campaign gives you a messy answer.
  3. Write down your hypothesis. One sentence: “Switching to Maximise Conversions will lower our cost per lead.” If you cannot write it, you do not have a test — you have a fiddle.
  4. Create the experiment. Campaigns → Experiments → the plus button → Custom experiment. Google clones the base campaign into a draft.
  5. Change exactly one thing. In the draft, not the base. This is the rule that gets broken most.
  6. Set the split. 50-50 for a straight comparison, unless you have reason to be cautious (next section).
  7. Set the dates and the goal metric. Four weeks minimum, and pick the metric that matches the hypothesis — conversions or cost per conversion, not clicks.
  8. Leave it alone. No mid-test edits to either arm. Every change resets what you are measuring.
Key takeaway: One variable, one hypothesis, four weeks, no touching. A test you interfere with is not a test.

5. 50-50, or Something Safer?

Quick Answer: Use 50-50 almost always. A cautious 80-20 split feels safer but starves the variant of data, so the test runs far longer — and a long test on a small account is the one thing you cannot afford. Protect yourself with test duration and a clear hypothesis, not with a lopsided split.

The instinct to give the experiment only 20% of traffic is understandable. It is also self-defeating: the variant now needs several times as long to earn a verdict, and on a modest budget it never earns one.

The honest trade-off: a 50-50 split does risk half your traffic on an unproven version for a month. Accept that, or do not test. No third option is fast, safe and readable at once.

Want the test designed before you spend the month?

ZenWeb sizes the test against your actual conversion volume first. Get a Google Ads account review →


6. Reading the Result Without Fooling Yourself

Quick Answer: The arm leading at day 7 is only a coin flip better than chance — it still leads at day 45 barely four times in ten. By day 28 that rises to nearly nine in ten. Early leads are noise wearing a costume, and the temptation to act on them is the single most expensive habit in Google Ads experiments.

Google will happily show you a winner in week one. It shows what the data says so far; it does not promise the data has settled. Here is how often the early leader still led at the end.

How Often the Leading Arm Was Still Leading at Day 45
Share of Google Ads custom experiments in which the arm leading on a given day was still the leading arm on day 45, measured at days 7, 14, 21, 28 and 35, across ZenWeb-managed Malaysian SME campaigns 2024 to 2026.
If you called it on…Still the winner at day 45Verdict
Day 741%Worse than a guess
Day 1458%Still a coin flip
Day 2174%Suggestive, not settled
Day 2886%Safe to act on
Day 3594%Settled

Source: Aggregated from ZenWeb-managed Google Ads campaigns, Malaysia, 2024–2026. Custom experiments that ran a full 45 days without mid-test edits.

Three habits keep the reading honest:

  • Judge on the goal metric only. The one you named in the hypothesis. Hunting through every column until one favours your preferred arm is not analysis — it is shopping.
  • Check the conversions are the same kind of conversion. If one arm’s “wins” are newsletter signups and the other’s are quote requests, the winner is an illusion. This is where your attribution model quietly decides the result for you.
  • Accept “no difference” as a real answer. It means the change does not matter — useful, and it frees the slot for a better test. Our guide to inconclusive A/B test results covers what to do next.

The discipline travels beyond Google Ads experiments, too. If your team tests across channels, the ground rules in A/B testing for marketers are worth ten minutes.

Key takeaway: A day-7 winner is a rumour. A day-28 winner is a result. The gap is where most testing budgets die.

7. Who Actually Runs Experiments — And Who Should

Quick Answer: Experiment use rises sharply with spend — and so does the quality of the result. Below RM 3,000 a month, few accounts test, only about three in ten tests conclude, and nearly a third of the “winners” get rolled back within a quarter. That rollback rate is the clearest sign a test was read too early.

Experiment Use and Outcomes by Monthly Google Ads Spend, Malaysian SMEs
Share of accounts running at least one experiment per quarter, share of experiments that reached a clear verdict, and share of applied winners later rolled back, grouped by monthly Google Ads spend band, across ZenWeb-managed Malaysian SME accounts 2024 to 2026.
Monthly spendRun ≥1 test a quarterTests reaching a verdictWinners rolled back
Under RM 3,00012%29%31%
RM 3,000–7,99934%51%19%
RM 8,000–19,99961%68%11%
RM 20,000 or more88%79%6%

Source: Aggregated from ZenWeb-managed Google Ads campaigns, Malaysia, 2024–2026. “Rolled back” = an applied experiment winner reversed within one quarter.

The last column is the one to sit with. On the smallest accounts, almost a third of applied winners get undone — the test said one thing, the next quarter said another. That is not bad luck. It is what happens when a verdict is squeezed from too thin a sample.

If your account sits in the bottom band, the highest-return work is not running Google Ads experiments. It is the plumbing: fixing tracking, tightening where your ads actually run, and checking your auction insights to see who you are bidding against. Those are gains you can bank without spending a month proving them.

Key takeaway: Experiments are a tool for accounts with volume. Below RM 3,000 a month, the test will usually tell you something that turns out not to be true.

8. When Not to Run an Experiment

Quick Answer: Skip the test when the change is obviously right, when the account is too small to read it, when a festive peak sits in the middle of it, or when the thing you want to test is not a variable at all — like chasing a higher optimisation score, which is Google’s opinion of your account, not a lever.

Not everything deserves a month of proof. Skip the experiment when:

  • The change is obviously correct. Broken tracking, an ad pointing at a dead page, missing negative keywords. Just fix them.
  • Volume is too thin. Under fifteen conversions a month in the base campaign, the result will not hold.
  • A festive peak lands mid-test. Either wait, or extend the test well past the peak on both sides.
  • You are chasing a score rather than a result. Google’s recommendations are suggestions, not findings — see whether you should chase a 100% optimisation score.
  • You cannot leave it alone. If someone will edit the campaign mid-test, the test is already dead.
Key takeaway: Google Ads experiments are expensive in time and traffic. Spend them on the questions you genuinely cannot answer any other way.

9. The Verdict

Quick Answer: Google Ads experiments are the only honest way to settle a big question in your account — but only if you have the volume to read them and the patience not to peek. Test bidding and landing pages, split 50-50, run four weeks, judge one metric, and treat “no difference” as a real answer.

Most advice on Google Ads experiments stops at the setup screen, as if the difficulty were finding the button. It is not. The difficulty is honesty — about whether your account is big enough to be tested, and whether the number in front of you has actually settled.

Get those two right and experiments become the account’s most valuable habit: a way to replace opinions with evidence, one question at a time. Get them wrong and you spend a month buying a result that reverses next quarter. The same discipline pays off on your Quality Score and your cost per lead — and if you want the whole account run this way, that is what ZenWeb’s Google Ads management is for.


10. Frequently Asked Questions

How long should a Google Ads experiment run?

Four weeks at an absolute minimum, and longer on low-volume accounts. In ZenWeb-managed campaigns, a base campaign with 30–59 conversions a month took around 38 days to produce a stable result; under 15 conversions a month, most experiments never produced one at all.

Do Google Ads experiments cost extra money?

No. The experiment shares the base campaign’s budget rather than adding to it. A 50-50 split sends half your existing budget to each arm, so total spend stays the same — but half of it is now buying information as well as clicks.

Can I run an experiment on a Performance Max campaign?

Not as a straight A/B split the way you can with Search. Performance Max supports uplift-style experiments that measure what PMax adds on top of your other campaigns. For a clean two-arm comparison, use a Search campaign.

What happens when the experiment ends?

You choose: apply the variant to your base campaign, keep it as a new standalone campaign, or end it and discard the changes. Nothing applies automatically — leave it, and the base campaign carries on as before.

Can I change an experiment while it is running?

You can, and you should not. Any mid-test edit resets what you are measuring, so the days before the edit no longer compare cleanly with the days after. If a change is truly urgent, end the test, make the change, and start fresh.

Tired of arguing about what to change next?

ZenWeb is a Google Partner managing Google Ads for 500+ Malaysian businesses. We will tell you whether your account has the volume to be tested — and if it does, design the experiment, run it properly, and read it honestly.

Talk to ZenWeb about your Google Ads

Table of Contents

Table of Contents

See Also

Long Tail Keywords: Easier Rankings, Better Sales Leads

Long Tail Keywords: Easier Rankings, Better Sales Leads

WooCommerce SEO: Rank Your WordPress Store in 2026

WooCommerce SEO: Rank Your WordPress Store in 2026

TikTok Ads Malaysia: What Actually Works for SMEs 2026

TikTok Ads Malaysia: What Actually Works for SMEs 2026

Get A Free Proposal

Complete the form and our team will contact you to discuss your goals. Let’s grow your business.

Meowketing Specialist

Online

Today

Meow! 👋

We are Official Google Partner,
Ask us anything about Marketing!