AdCopilotby Atromx

Ad Copy Testing With AI: A Framework That Holds Up

RSAs killed the clean A/B split. The honest replacement is asset-level analysis plus deliberate variants — and AI does the tedious half of both.

Updated 2026-08-10Atromx IntelligenceGoogle Ads · Search, PMax, Display, YouTube, Demand Gen
The short answer

Ad copy testing in the RSA era means two disciplines, not one. Asset-level analysis prunes weak headlines using performance labels and impression data — an AI reads and ranks these across the account in one prompt. Deliberate variant tests answer real hypotheses: launch one changed message per ad group, run it for weeks, and judge on conversion metrics, not impressions. Classic A/B splits died when Google started assembling ad combinations per auction.

Ad copy testing used to be tidy: two ads, even rotation, wait, keep the winner. Responsive search ads ended that — Google now assembles combinations from your assets per auction, so the unit you wrote is no longer the unit that serves. The honest replacement is a two-part framework: asset-level pruning to remove weak material, and deliberate variant tests to answer real questions. An AI connected to the account does the tedious half of both — the reading, ranking and tracking — leaving you the hypotheses and the verdicts.

Why old-school ad splits died with RSAs

Three mechanics killed the classic split. First, RSAs are combinatorial: with up to fifteen headlines mixed per auction, "ad A versus ad B" hides thousands of actual creative variations. Second, serving is optimised by default — Google shows the combinations it predicts will work, so exposure is never evenly divided and impression counts reflect Google's preferences as much as user response. Third, conversion data attaches to the ad, not to the individual headline, so the question "which headline converted" has no recorded answer.

Pretending otherwise produces confident nonsense. The framework that holds up accepts the machinery: prune assets on the evidence that exists, and test messages at a level where conversions genuinely are recorded. The machinery is not going back, either — combinatorial ads and optimised serving are the platform's direction everywhere, which makes learning to test inside them the durable skill rather than the workaround.

What asset-level reporting can and cannot tell you

Google labels RSA assets — ratings like Low, Good and Best, plus a learning state — and reports how often each was served. A connected model turns that into an account-wide view in one prompt:

Across all RSAs in account 123-456-7890, list assets by performance
label and serving share. Which headlines are rated lowest, what do
the strongest-serving headlines have in common, and which ad groups
have assets stuck in learning?

What this evidence supports: removing persistently low-rated headlines, spotting the patterns in what Google chooses to serve (lengths, angles, question forms), and finding ad groups where every asset says the same thing. What it does not support: declaring a conversion winner, or treating serving share as merit — an asset can serve heavily because Google predicted clicks, not because it sells. Labels are a pruning signal. Prune with them, and stop there.

Designing variant tests an AI can launch and track

A real test starts with a message hypothesis, not a formatting itch: leading with price beats leading with credentials; the guarantee matters more than the feature list. Then:

  1. One variable. Build a variant RSA that changes only the message under test — same keyword coverage, same pins, per the RSA drafting workflow.
  2. One arena. Run it in ad groups with enough conversion volume to eventually support a verdict. Low-volume ad groups produce eternal ties.
  3. Launch through the connector. The AI creates the variant — paused for review, then live with your approval — and logs the start date and the hypothesis.
  4. Pre-commit the verdict rules. Metric (conversion rate or cost per conversion), judgement date, and the difference big enough to act on. Write these down before launch; they are the antidote to week-two enthusiasm.

For wording-level hypotheses that span many ad groups at once — a claim swap, a price mention — Google's ad variations feature is worth knowing: it applies a find-and-replace style change across matching RSAs and splits traffic against the original. Its limit is its shape: it tests a text substitution, not a message architecture, so the deliberate-variant workflow above remains the tool for bigger swings. And for strategy-level changes — bidding, landing pages, structures — the platform's own split-traffic machinery is the better instrument; that is covered in Google Ads experiments with AI.

Reading results without fooling yourself

Impressions are not verdicts. The three self-deceptions that account for most false wins:

  • Serving share read as preference. Google showed variant B more, so B "won". Serving share is Google's prediction, not the market's answer — judge on conversion metrics only.
  • The early call. Leads in week one are usually noise plus learning-phase artefacts. The judgement date exists precisely because the first fortnight lies.
  • Metric shopping. If conversion rate disappoints, CTR suddenly becomes the success metric. Pre-committing the metric closes this door.

Put the AI on discipline duty: "Compare the test and control on the pre-committed metric, state whether the difference clears the threshold we set, and say plainly if the data cannot support a call yet." A model with the numbers in front of it is far more comfortable saying "not yet" than a human three weeks into wanting an answer. Ad strength, for the same reason, stays out of the verdict entirely — it grades inputs, not outcomes.

The sample-size intuition, without the theatre

You do not need a statistics degree; you need one reframe: tests are measured in conversions per arm, not in weeks. A difference between variants deserves trust only when each arm has accumulated enough conversions that the gap could not plausibly be luck — and small gaps need far more evidence than large ones. Ask the model to run the honest version of this arithmetic before launch: given current volume, how big a difference could this test even detect by the judgement date? If the answer is "only a huge one", the right response is a bolder variant or a bigger arena — decided up front, not discovered in the post-mortem.

A quarterly copy-testing calendar run through prompts

Testing that is not scheduled does not happen. A quarter's rhythm, one saved prompt each:

  • Month 1 — prune. Run the asset-label sweep. Replace persistently low-rated headlines with fresh drafts grounded in current search terms. This is maintenance, not testing, but it clears the ground.
  • Month 2 — launch. One message hypothesis, in the two or three highest-volume ad groups. Log hypothesis, start date, verdict rules.
  • Month 3 — judge and record. Run the verdict prompt on the judgement date. Win, loss or no-call, write one line about what was learned and roll winners into the standing copy.

One genuine test per quarter, judged honestly, compounds faster than a dozen splits eyeballed weekly — because its conclusions survive. Keep the log alive across quarters too — hypotheses that lost, angles already mined, the messages that won and when — because the second most expensive test is the one that quietly reruns last year's loser. The model's contribution is not creativity alone; it is that the boring parts — pulling labels, tracking dates, holding you to your own rules — now cost a sentence each.

Frequently asked questions

Does Google still support ad rotation settings for testing?

Rotation settings still exist at campaign level, but the default is optimised serving — Google prefers the ads and combinations it predicts will perform, and with responsive search ads the assembly happens inside the ad anyway. Even rotation cannot give you a clean split of RSA assets. Use rotation preferences only when running deliberate multi-ad tests, and accept that optimised serving is the house style now.

Can the AI tell me which headline won?

It can rank asset performance labels and impression counts, and tell you which headlines Google chose to serve most — that is what asset reporting exposes. What nobody can tell you is per-headline conversion attribution inside an RSA, because conversions attach to the ad, not the asset. Treat labels as serving signals, prune the clear losers, and reserve the word won for variant tests judged on conversions.

How long should a copy variant test run?

Until the conversion numbers can support a decision — for most accounts that is four to six weeks, longer for low-volume ad groups. Ending a test the first week one variant looks ahead is the classic self-deception, because early leads are mostly noise and serving-share artefacts. Decide the judgement date and the metric before launch, and let the AI hold you to both.

The offer

Try it on your own account for a week

The full set of tools for the week, so you can see what it actually does — and it still cannot delete anything. No cost, no card, no contract: you connect your own Google account and can withdraw the access whenever you like.

  • Up to 5 accounts
  • One week
  • Full tools
  • No card
Keep reading