Geo Holdout Tests a Sub-$1M DTC Brand Can Run
Your dashboard says 4.2 ROAS. Your bank account disagrees. The gap is the question every operator eventually has to answer: how many of those reported conversions would have happened anyway, with no ad at all? Attribution platforms credit clicks; they don’t run the counterfactual. A geo holdout test does — and it is not enterprise-only equipment. With two matched regions and a spreadsheet, a brand spending well under seven figures a year can measure true incremental lift in about four weeks.
For the neighboring economics, compare Guest Checkout vs Forced Sign-Up: Run the Conversion Math and use Reporting Window vs Optimization Window, Honestly to validate the measurement decision.
Why platform ROAS can’t answer the question
Conversion reporting is correlational. It tells you that someone saw or clicked an ad and later bought. It cannot tell you whether the ad caused the purchase. A meaningful share of conversions credited to prospecting and almost all conversions credited to brand-term retargeting are people who were already going to buy. The platform happily takes the credit.
The only way to separate caused-sales from would-have-happened-anyway is a controlled experiment: withhold advertising from one comparable population, keep it running for another, and measure the difference. That difference is incrementality. Geo testing is the least expensive, most robust way for a small brand to do it, because geography is something you can actually control on the platform — and it doesn’t require user-level tracking, which keeps getting harder.
Why a geo holdout test fits a small ecommerce brand
Two reasons people assume this is out of reach, and why both are wrong:
- “I don’t have enough volume.” You don’t need millions of impressions. You need enough conversions in the holdout cell to detect a difference. The lever is the split: hold out 20–35% of your footprint rather than 5%, so the test cell carries enough events to see a real gap.
- “I need a fancy tool.” The math is a difference-of-differences. A spreadsheet handles it. Paid geo-experiment platforms add automation and confidence intervals, but the core logic is something you can run by hand and trust because you can see every number.
The trade-off is honest: you give up some revenue in the holdout region for the test window. That’s the cost of buying the truth. Size the holdout so the foregone spend is affordable — this is a measurement investment, not a growth campaign.
What you need before you start
- A geographically splittable footprint. Your sales need to come from more than one area you can target separately on the platform. If 90% of revenue is one metro, geo testing won’t work — use a different method.
- Clean regional revenue data. Pull orders by shipping or billing region from your store backend, not from the ad platform. The platform’s regional numbers are the thing you’re auditing; don’t grade your own homework.
- A stable baseline. Four to six weeks of pre-test history per region, ideally with no major promotion, stockout, or price change mid-flight.
- Patience for the window. Run the test 3–4 weeks minimum so the holdout region’s demand has time to actually soften once ads stop. Pausing ads on Monday and reading results on Friday measures nothing.
Step by step
1. Build matched cells. Don’t pick one region for test and one for control at random. Group your geographies into a test cell (ads stay on) and a holdout cell (ads go dark) so the two cells have similar baseline weekly revenue and similar week-to-week trend shape. Matching on trend matters more than matching on size — you’ll normalize for size in the math.
2. Lock the baseline ratio. Over your pre-test weeks, compute the holdout cell’s revenue as a ratio of the test cell’s revenue. Say the holdout historically runs at 0.40 of the test cell. That ratio is your prediction engine.
3. Go dark cleanly in the holdout. Pause — don’t just lower budget on — all paid delivery into the holdout geographies. Watch for leakage: broad campaigns, lookalikes, and Advantage+ placements can spill outside your intended targeting. Confirm in delivery reporting that holdout regions are actually receiving near-zero impressions. Keep the test cell running exactly as before; change nothing else.
4. Hold everything else still. No new promo, no price change, no creative refresh, no email blast that hits one cell harder than the other. Any of these will contaminate the read.
5. Measure the gap. Each test week, predict what the holdout should have done if ads were still working: multiply the test cell’s actual revenue by your baseline ratio. Compare to what the holdout actually did with ads off. The shortfall is your incremental revenue.
Reading the result
Here’s the worked logic with round numbers. In a test week:
- Test cell (ads on) does 100 in revenue.
- Baseline ratio is 0.40, so the holdout should do 40 if ads still mattered there.
- Holdout (ads off) actually does 34.
The 6-unit shortfall is revenue the ads were producing — roughly 15% of the holdout’s expected revenue was incremental to advertising. Now compare that to what the platform claimed it was driving in that same region. If the platform reported the holdout’s ad-driven revenue at, say, 12 units a week and the true incremental was 6, your incrementality factor is about 0.5 — the platform is roughly double-counting.
Apply that factor across the account and your real blended return drops accordingly. A reported 4.0 ROAS at a 0.5 incrementality factor is a true incremental ROAS near 2.0. Whether that’s good depends entirely on your contribution margin: if your margin after COGS, shipping, and fulfillment doesn’t cover a 2.0, you’re buying revenue at a loss and the dashboard is hiding it.
Treat every benchmark here — the 15% incremental share, the 0.5 factor, even “~50 conversions to read a cell” — as illustrative planning ranges to pressure-test against your own data, not fixed truths. Your numbers will differ, and the whole point is to learn yours.
Where these tests go wrong
- Holdout too small. A 5% holdout has too few conversions to distinguish signal from noise. Bias toward a larger holdout even though it costs more.
- Spillover. People travel, ship to gifts, and use shared networks. Some ad exposure crosses your geographic line. It dampens measured lift; read a clear result as a conservative floor, not an exact point.
- Reading too early. Demand has inertia. Cutting ads doesn’t zero out sales the next day — it decays. Short windows understate the effect.
- Confounds. A competitor’s promo, a stockout, a viral moment in one cell — any of these break the match. Keep an eye on the test cell’s trend; if it does something weird, the read is suspect.
- One test, forever. Incrementality drifts with creative fatigue, audience saturation, and seasonality. A read is a snapshot, not a constant. Re-run quarterly.
The takeaway
You don’t need a data science team to stop flying blind. Pick two matched cells, hold out a meaningful slice, go dark cleanly for a month, and compare actual-to-predicted in a spreadsheet. The output — your true incrementality factor — is the multiplier that turns a flattering platform ROAS into a number you can hold against margin and actually trust.
This is exactly the kind of analysis Bach is built to keep honest: it reads your delivery and store data and surfaces where reported return and incremental return diverge, so you’re scaling on caused revenue, not credited revenue. But you can prove the principle to yourself first, this week, with nothing more than the data you already have.