Skip to content
Bach.ai

The Honesty Test for AI Creative: Incremental or Harvesting?

Your new AI-generated creative posted a 4.2 ROAS in its first week. The dashboard crowned it the winner, the budget shifted, and everyone moved on. Here’s the uncomfortable question almost nobody asks: did that creative actually create demand, or did it just take credit for buyers who were already going to purchase? Platform-reported ROAS cannot tell the difference, and the gap between “looks like a winner” and “is a winner” is where most creative budgets quietly leak.

For the adjacent tooling decision, compare Localizing Ad Creative at Scale With AI: Beyond Translation and use Auto-Tagging Ad Creative: Read What Truly Drives Sales to evaluate the operating trade-off.

Why platform ROAS lies to you about creative

The optimizer does its job too well. When you launch a new ad, the delivery system finds the people plausibly to convert and serves it to them. Many of those people were already deep in the consideration set — they’d seen you before, they were ready, and they’d have bought through some other touchpoint anyway. The new creative shows up in the conversion path, gets the attribution credit, and your dashboard reports a high ROAS.

This is harvesting: capturing conversions that would have happened without the ad. The opposite is incrementality: conversions that exist because the ad ran. Platform ROAS counts both as wins. It has no way to separate them, because it never observes the counterfactual — what would have happened if the ad hadn’t been served.

AI creative makes this trap sharper, not softer. You can now generate twenty variants in an afternoon, and the platform will reliably surface the one with the highest reported ROAS. But “highest reported ROAS” disproportionately selects for the variant that harvested most efficiently, not the one that expanded demand. You can iterate yourself straight into a portfolio of beautiful, well-converting ads that move almost no new revenue. The optimization loop feels productive while the contribution-margin needle stays flat.

What an honest read actually requires

The only verdict that survives scrutiny is a measured counterfactual. That means running an incrementality test — deliberately withholding the creative from a comparable slice of demand and measuring the difference in outcomes, not the absolute performance of the exposed group.

Two methods carry the weight here.

Holdout (conversion lift) tests. A randomly selected portion of your addressable audience is held out from seeing the new creative. After the test window, you compare conversions in the exposed group versus the held-out group. The delta is your lift. If the new creative drove 4.2 reported ROAS but the holdout converted nearly as well on its own, your incremental ROAS is a fraction of the headline number — sometimes close to zero.

Matched-cell splits. Where audience-level holdouts aren’t clean, split your addressable demand into matched test and control cells — paired on baseline trend, seasonality, and revenue scale (matched on behavior, never chosen for any reason but comparability). Run the creative into the test cells, keep the control cells on your existing rotation, and read the difference in total sales — not platform-attributed sales, total sales, because that’s the only number the optimizer can’t reshuffle in its own favor.

Both methods share the same discipline: you are measuring against a world where the creative didn’t run. That’s the honesty test. If a creative can’t beat the absence of itself, it isn’t a winner no matter what the attribution column says.

Measure margin, not the anointed metric

Even incremental ROAS is the wrong endpoint. ROAS is revenue over spend, and revenue is not what your business keeps. A creative can be genuinely incremental and still lose money if it pulls in discount-hunters, low-AOV first orders, or buyers who never repeat.

Read incrementality on contribution margin instead — revenue minus COGS, minus shipping and fulfillment, minus the ad cost itself. A creative that drives incremental sales at a CPA above your margin is a real machine for losing money faster. The honest scorecard is one line:

Did total contribution margin go up in the test cell versus control, by more than the cost of running the test creative?

If yes, you have an incremental winner worth scaling. If no, you have a harvester — possibly an efficient one, but it earns its budget by reshuffling existing demand, not by growing it.

A simple way to hold the two readings side by side:

Signal Platform dashboard Incremental + margin read
What it measures Attributed conversions on exposed users Lift vs. a no-ad counterfactual
Can it be gamed by the optimizer Yes — favors harvesting No — total outcomes only
Endpoint Revenue / spend Contribution margin delta
Verdict on a “4.2 ROAS” ad Winner Unknown until the holdout reports

A practical protocol you can run

You don’t need a measurement science team to do this credibly. You need discipline about three things: a real counterfactual, enough signal, and patience.

  1. Pick one creative decision that matters. Not twenty micro-variants — one genuine bet (a new concept, a new format, an AI-generated angle you’re tempted to scale). Incrementality testing is expensive in attention; spend it on decisions with real budget consequences.
  2. Build a clean counterfactual. Either a randomized audience holdout or a matched-cell split. The control must be genuinely comparable on baseline trend and revenue, or the read is noise dressed as insight.
  3. Size the test for signal, then wait. Lift tests need meaningful conversion volume in both cells before the difference stabilizes — think in terms of enough recent conversions per cell to move past noise, not a fixed magic number, and let the optimizer exit its learning period before you start counting. Reading a holdout after three days is reading randomness.
  4. Score on contribution margin. Pull total sales (not attributed) for both cells, subtract COGS, fulfillment, and the test spend, and compare. That delta is your verdict.
  5. Make the call, then re-test cadence. Incrementality decays. A creative that was incremental at one spend level can flip to harvesting as you scale it into the same audience. Re-read periodically; don’t treat one green test as permanent.

This is exactly the kind of read a tool should surface for you instead of leaving you to reverse-engineer it from a dashboard. Bach is built to flag when a “winning” creative’s lift and margin story don’t match its reported ROAS — and, being read-only until you approve, it shows you the honest verdict before anything changes in the account.

The takeaway

Treat every “winning” AI creative as a suspect until it passes the honesty test. The dashboard’s anointed winner is a hypothesis, not a result. The result lives in the gap between an exposed cell and a held-out one, measured on contribution margin, after the optimizer has had enough signal to settle. Run that read on the few creative decisions that actually move budget, and you’ll stop paying premium rates to harvest demand you already owned — and start funding the creatives that genuinely make the pie bigger.

See what your Meta ads are really costing you.

Connect your account and Bach ranks every revenue leak in minutes — each with the money it costs and a one-tap fix. Free for 7 days, no credit card.

Start Free Audit
Start your free audit