Ad copy testing used to be tidy: two ads, even rotation, wait, keep the winner. Responsive search ads ended that — Google now assembles combinations from your assets per auction, so the unit you wrote is no longer the unit that serves. The honest replacement is a two-part framework: asset-level pruning to remove weak material, and deliberate variant tests to answer real questions. An AI connected to the account does the tedious half of both — the reading, ranking and tracking — leaving you the hypotheses and the verdicts.
Why old-school ad splits died with RSAs
Three mechanics killed the classic split. First, RSAs are combinatorial: with up to fifteen headlines mixed per auction, "ad A versus ad B" hides thousands of actual creative variations. Second, serving is optimised by default — Google shows the combinations it predicts will work, so exposure is never evenly divided and impression counts reflect Google's preferences as much as user response. Third, conversion data attaches to the ad, not to the individual headline, so the question "which headline converted" has no recorded answer.
Pretending otherwise produces confident nonsense. The framework that holds up accepts the machinery: prune assets on the evidence that exists, and test messages at a level where conversions genuinely are recorded. The machinery is not going back, either — combinatorial ads and optimised serving are the platform's direction everywhere, which makes learning to test inside them the durable skill rather than the workaround.
What asset-level reporting can and cannot tell you
Google labels RSA assets — ratings like Low, Good and Best, plus a learning state — and reports how often each was served. A connected model turns that into an account-wide view in one prompt:
Across all RSAs in account 123-456-7890, list assets by performance
label and serving share. Which headlines are rated lowest, what do
the strongest-serving headlines have in common, and which ad groups
have assets stuck in learning?
What this evidence supports: removing persistently low-rated headlines, spotting the patterns in what Google chooses to serve (lengths, angles, question forms), and finding ad groups where every asset says the same thing. What it does not support: declaring a conversion winner, or treating serving share as merit — an asset can serve heavily because Google predicted clicks, not because it sells. Labels are a pruning signal. Prune with them, and stop there.
Designing variant tests an AI can launch and track
A real test starts with a message hypothesis, not a formatting itch: leading with price beats leading with credentials; the guarantee matters more than the feature list. Then:
- One variable. Build a variant RSA that changes only the message under test — same keyword coverage, same pins, per the RSA drafting workflow.
- One arena. Run it in ad groups with enough conversion volume to eventually support a verdict. Low-volume ad groups produce eternal ties.
- Launch through the connector. The AI creates the variant — paused for review, then live with your approval — and logs the start date and the hypothesis.
- Pre-commit the verdict rules. Metric (conversion rate or cost per conversion), judgement date, and the difference big enough to act on. Write these down before launch; they are the antidote to week-two enthusiasm.
For wording-level hypotheses that span many ad groups at once — a claim swap, a price mention — Google's ad variations feature is worth knowing: it applies a find-and-replace style change across matching RSAs and splits traffic against the original. Its limit is its shape: it tests a text substitution, not a message architecture, so the deliberate-variant workflow above remains the tool for bigger swings. And for strategy-level changes — bidding, landing pages, structures — the platform's own split-traffic machinery is the better instrument; that is covered in Google Ads experiments with AI.
Reading results without fooling yourself
Impressions are not verdicts. The three self-deceptions that account for most false wins:
- Serving share read as preference. Google showed variant B more, so B "won". Serving share is Google's prediction, not the market's answer — judge on conversion metrics only.
- The early call. Leads in week one are usually noise plus learning-phase artefacts. The judgement date exists precisely because the first fortnight lies.
- Metric shopping. If conversion rate disappoints, CTR suddenly becomes the success metric. Pre-committing the metric closes this door.
Put the AI on discipline duty: "Compare the test and control on the pre-committed metric, state whether the difference clears the threshold we set, and say plainly if the data cannot support a call yet." A model with the numbers in front of it is far more comfortable saying "not yet" than a human three weeks into wanting an answer. Ad strength, for the same reason, stays out of the verdict entirely — it grades inputs, not outcomes.
The sample-size intuition, without the theatre
You do not need a statistics degree; you need one reframe: tests are measured in conversions per arm, not in weeks. A difference between variants deserves trust only when each arm has accumulated enough conversions that the gap could not plausibly be luck — and small gaps need far more evidence than large ones. Ask the model to run the honest version of this arithmetic before launch: given current volume, how big a difference could this test even detect by the judgement date? If the answer is "only a huge one", the right response is a bolder variant or a bigger arena — decided up front, not discovered in the post-mortem.
A quarterly copy-testing calendar run through prompts
Testing that is not scheduled does not happen. A quarter's rhythm, one saved prompt each:
- Month 1 — prune. Run the asset-label sweep. Replace persistently low-rated headlines with fresh drafts grounded in current search terms. This is maintenance, not testing, but it clears the ground.
- Month 2 — launch. One message hypothesis, in the two or three highest-volume ad groups. Log hypothesis, start date, verdict rules.
- Month 3 — judge and record. Run the verdict prompt on the judgement date. Win, loss or no-call, write one line about what was learned and roll winners into the standing copy.
One genuine test per quarter, judged honestly, compounds faster than a dozen splits eyeballed weekly — because its conclusions survive. Keep the log alive across quarters too — hypotheses that lost, angles already mined, the messages that won and when — because the second most expensive test is the one that quietly reruns last year's loser. The model's contribution is not creativity alone; it is that the boring parts — pulling labels, tracking dates, holding you to your own rules — now cost a sentence each.