Google Ads has a built-in A/B testing tool, and most accounts have never run it once. That is not because advertisers dislike evidence — it is because an experiment costs three chores: designing a clean test, building the variant, and reading the result honestly. A connected AI agent takes all three off your hands in the time it takes to describe them — it designs the split, drafts the variant, and reads both arms honestly. What remains is the decision, which was always the valuable part.
Why experiment adoption is embarrassingly low
An experiment asks you to hold a question steady for a month. Accounts are run by people whose week is interrupts — budgets, disapprovals, a client asking why Tuesday was slow. Against that backdrop, the experiment workflow (pick a hypothesis, build a mirrored change, set a split, wait, then do arithmetic on two arms) keeps losing to "just change it and watch".
The cost of "change it and watch" is that you never learn anything cleanly. Performance moved — was it your change, seasonality, a competitor, or conversion lag? Without a control arm running alongside, every readout is a story. Experiments exist to replace stories with differences. The chores were the price, and the chores are now delegable.
What is worth testing (and what never reaches significance)
The brutal filter for test ideas is volume: an experiment can only detect a difference bigger than its noise. A useful triage, which a connected AI can run against your own numbers before you commit:
| Candidate test | Signal it needs | Verdict for most accounts |
|---|---|---|
| Bid strategy change (e.g. manual to target CPA) | Conversions per arm, weekly | Worth it if the campaign converts steadily |
| Target adjustment size (raise target CPA) | Conversions per arm | Worth it on your biggest campaign only |
| Landing page swap | Conversions per arm | Worth it — differences are often large |
| Broad match adoption on a tight campaign | Clicks and conversions | Worth it, fenced, with query review |
| One headline tweak in one RSA | Impressions per asset | Almost never — the ad shuffles assets anyway |
| Tiny budget move on a small campaign | Conversions | Never reaches a verdict — skip the ceremony |
The pattern: structural changes with campaign-level effects test well; small creative changes inside asset-shuffled ads mostly do not, because the auction never shows the same combination twice. Ask the agent directly: "given this campaign's weekly conversions, what is the smallest effect a six-week even-split test could plausibly detect?" If the honest answer is "only a huge one", test something bigger.
Setting up a custom experiment through conversation
The design conversation takes minutes because the AI has the account open in front of it. A real sequence looks like:
You: I want to test target CPA against maximise conversions on the Plumbing — Search campaign. Design the experiment.
Agent: Base campaign converts steadily on weekdays. Proposal: even traffic split, six weeks starting Monday, one change only — bid strategy on the trial arm. Success metric: cost per conversion; guardrail: conversion volume must not fall materially. Decision rule: adopt if the trial arm is clearly cheaper per conversion at similar volume; otherwise keep the base.
You create the experiment in the Google Ads interface — that is where the traffic-split machinery lives — pasting the specification the agent wrote. From then on the connector reads both arms like any other campaign, so monitoring is a prompt, not a pivot table. The agent's role is analyst and scribe, and every account change it proposes along the way passes through the approval loop like any other write.
Three creation-screen settings deserve their moment of attention, because they silently shape the readout. The split basis — whether users or searches are divided between arms — decides how clean the comparison is for repeat searchers. Sync settings decide whether edits to the base campaign copy across to the trial arm mid-test; leaving them on preserves the one-change discipline when someone inevitably touches the base. And the end date should match the decision date you pre-registered, so the test cannot quietly run forever while everyone waits for it to feel conclusive.
One discipline the conversation should enforce: one change per experiment. The moment a trial arm differs in two ways, the readout can attribute the result to neither. Agents are actually better at this discipline than people, because they will flag the second change instead of sneaking it in.
Reading results without p-value theatre
The failure mode of DIY experiments is not bad arithmetic — it is motivated reading. Three rules keep the readout honest, and all three are easy to delegate:
- Pre-register the decision. The adopt/reject rule was written on day one, before data existed. The agent's weekly report restates it every time, which makes quiet goalpost-moving embarrassing.
- Respect conversion lag. The last days of a test are always missing late conversions. Ask the agent to read arms with a lag-aware cutoff — compare periods both arms have had time to complete.
- Impressions are not verdicts. An arm can win attention and lose economics. The readout ranks the pre-chosen success metric first, volume guardrail second, and everything else as commentary.
A weekly prompt that does all three: "Report the experiment: both arms on the success metric and guardrail, excluding the lag window, and restate the decision rule. Do not recommend early adoption."
Rolling out winners and logging what you learned
When the decision date arrives, the call is usually undramatic — one arm is better or the difference is noise, and the pre-registered rule says what to do with each outcome. Apply the winner, then have the agent write the two-paragraph log entry: hypothesis, result, decision, and what to test next. Accounts accumulate these into something rare in PPC: institutional memory that survives staff changes.
Then start the next one. The point of removing the chores is not to run one clean test — it is to make testing the default way changes enter the account. One live experiment at any time is a modest, sustainable pace that most accounts have never once managed. And log the losers with the same care as the winners: a test that proved the flashy bid strategy does nothing for your account just saved every future month it would otherwise have been re-proposed in.
If your account has never run a test, the first conversation is free to have: connect it and ask what is worth testing given your volume. Start a free pilot and run the triage prompt from this page.
Frequently asked questions
How long should a Google Ads experiment run?
Until the volume justifies a call, which for most accounts means four to six weeks — long enough to cover conversion lag and at least a few full weekly cycles. Ending early because one arm is ahead is the classic sin: early leads are usually noise. Set the decision date when you launch the test, and let the AI report weekly without acting until that date arrives.
Can experiments test bid strategy changes safely?
Yes — that is their best use. A bid strategy change applied directly to a live campaign risks the whole campaign's performance during learning; an experiment confines the change to a slice of traffic while the original keeps running unchanged alongside it. If the new strategy struggles, the damage is fenced to the test arm, and ending the experiment restores full traffic to the proven setup.
Does the AI launch the experiment itself?
The agent does the design and the reading; the experiment object itself is created in the Google Ads interface, where the traffic-split controls live. In practice the AI hands you a complete specification — base campaign, the one change, split, dates, success metric and decision rule — and once the test is live it reads both arms through the connector and reports the gap with its confidence caveats attached.
Try it on your own account for a week
The full set of tools for the week, so you can see what it actually does — and it still cannot delete anything. No cost, no card, no contract: you connect your own Google account and can withdraw the access whenever you like.
- Up to 5 accounts
- One week
- Full tools
- No card
- Autonomous agentsLevels of autonomy in Google Ads management, which optimisation work is safe unattended versus which needs approval, and why irreversible actions should not be automated.
- Google Ads MCP serverWhat a Google Ads MCP server is, how free self-hosted servers compare to a hosted one, the full tool list AdCopilot exposes, and what you need to connect.
- Connect ClaudeStep-by-step instructions for adding a Google Ads MCP connector to Claude Desktop, claude.ai and Claude Code, including what to ask it first and how to revoke access.
- Connect ChatGPTStep-by-step instructions for adding a Google Ads MCP connector to ChatGPT, what it can read and change, and how to withdraw access.