Seasonality and Promos: Confounders That Fake Lift
Your last lift test probably lied to you. Not because the math was wrong, but because the baseline you measured against was moving on its own the entire time. A test compares a treated group to a control and calls the gap “incrementality” — but if demand, discounts, returns, and inventory were all shifting underneath both groups, the gap is measuring the calendar, not your ads. Most DIY tests never separate the two, so they ship a number that feels rigorous and is quietly fiction.
For the neighboring economics, compare Abandoned-Cart Flows That Recover Revenue Without Discounts and use Meta Conversion Lift: What It Proves and Hides to validate the measurement decision.
The baseline is not a flat line
The whole premise of an incrementality test is simple: what would have happened anyway? That counterfactual is the baseline. The problem is that operators treat the baseline as a constant — last month’s run-rate, a trailing average, a “normal week.” It is never constant. Organic demand drifts week to week. Your own promo calendar yanks conversion rate up and down. Returns claw revenue back weeks after the sale. Stockouts cap how much you could have sold no matter how good the ad was.
When the baseline moves more than the ads do, the test result is dominated by noise you mislabeled as signal. The reported lift is real arithmetic on top of a fake foundation. This is the heart of why incrementality test confounders matter: a confounder is anything that moves both the outcome and correlates with your test timing or test groups, and every one of the four below does exactly that.
The four confounders that fake lift
Seasonality. Demand has a natural rhythm — weekly cycles, monthly pay-cycle bumps, slow stretches and busy stretches that have nothing to do with your spend. If you run a two-week test that happens to start in a rising-demand window, the treated period looks inflated even with zero ad effect. Run the same test in a falling window and you’ll “prove” your ads do nothing.
Promos. A discount, a bundle, a free-shipping threshold, or a site-wide sale changes conversion rate, average order value, and buyer intent simultaneously. If a promo overlaps your test — even partially, even on just the treated cohort’s landing experience — you cannot tell whether the ad drove the purchase or the price did. Promos are the single common silent confounder because marketing and merchandising seldom sync calendars.
Returns. Top-line revenue is not contribution. A test that books revenue at checkout overstates lift in any category with a meaningful return rate, because returns land days or weeks later and disproportionately hit exactly the impulse-driven, discount-driven, new-customer orders that ads can generate. Measure gross and you flatter the ad. Measure net and the lift can shrink by a large fraction.
Stockouts. When a hero SKU goes out of stock, the ceiling on conversions drops independent of demand. If the stockout hits mid-test, you’ve capped the treated group’s possible outcome and your lift collapses — not because the ad stopped working, but because there was nothing to sell. Partial stockouts (one size, one variant) do the same thing in miniature and are even easier to miss.
How to control for each
You control confounders two ways: by design (so the confounder hits both groups equally) and by measurement (so you can see and net it out). Use both.
| Confounder | Design control | Measurement control |
|---|---|---|
| Seasonality | Randomize treatment/control concurrently; run a clean pre-period | Difference-in-differences vs. the pre-period trend |
| Promos | Freeze the promo calendar across the test window for both groups | Flag and segment any overlapping promo; exclude or model it |
| Returns | Use a settling window before reading results | Measure on returns-adjusted net revenue or contribution |
| Stockouts | Pre-check inventory depth for test SKUs | Log stock status daily; censor stocked-out days |
A few specifics worth internalizing:
- Randomize concurrently, not sequentially. The best-supported defense against seasonality and promos is that the control group lives through the exact same calendar as the treated group. A holdout — a randomized slice of your audience or a set of matched geographic units that sees no ads — absorbs every time-based shock equally. Before/after tests have no such defense and should be your last resort, not your default.
- Establish a pre-period. Measure both groups for a stretch before the test starts. If they were already drifting apart, you have a selection problem; if they tracked together, your difference-in-differences estimate is trustworthy. This single step exposes most baseline drift.
- Net out returns before you call it. Build a settling window into the test — long enough for the bulk of returns to post — and read results on contribution after returns, shipping, and discount, not on checkout revenue. The honest number is almost always smaller.
- Censor, don’t ignore, stockouts. Track inventory at the SKU and variant level daily. Days where a test SKU was unavailable can’t fairly count against the ad; drop them or model the cap explicitly rather than letting them silently deflate your lift.
Don’t undermine the test with the platform itself
Two delivery realities quietly corrupt homemade tests. First, the learning phase: a fresh test cell needs enough recent optimization-event signal before delivery stabilizes — think on the order of dozens of conversions per cell as an illustrative planning range, not a fixed rule, since Meta needs enough recent signal and the exact bar varies. Read results before the cell stabilizes and you’re measuring algorithmic warm-up, not ad effect. Second, attribution windows decide what counts as a converted exposure; a generous window inflates credit and a stingy one buries delayed conversions. Lock the window before the test and keep it identical across cells.
If you can use a properly randomized holdout — a true control that’s eligible but suppressed — lean on it. It’s the cleanest way to make seasonality, promos, returns, and stockouts hit treatment and control symmetrically, so the leftover gap is closer to genuine incrementality.
A quick worked example
Say a two-week test shows treated revenue running 30% above control, and you’re tempted to book a 30% lift. Now layer the confounders. A discount went live on day 4 for everyone, lifting baseline conversion ~10% — not your ad. The category returns roughly a fifth of new-customer orders, which won’t post until after checkout — shave the net. A best-seller variant was out of stock for three of the fourteen days, capping the treated upside. Adjust for all three and that headline 30% can land far lower in contribution terms. Same data, honest baseline — a completely different decision about whether to scale.
The takeaway
Before you trust any lift number, audit the baseline, not the ad. Run treatment and control concurrently with a randomized holdout, anchor to a clean pre-period, freeze the promo calendar, wait out returns, and censor stockout days. If you can’t control a confounder, at least flag and segment it so you know which direction it bends the result. This is exactly the kind of confounder-checking Bach AI does when it reads an account — it won’t call a result “lift” until the calendar, the discounts, the returns, and the shelf are accounted for. A smaller, honest number you can act on beats a big one that’s just the season wearing your ad’s name.