When MMM, MTA, and Incrementality Each Break Down
Every attribution debate eventually slams into the same wall: the method you trust most is wrong by more than your margin can absorb. On a fat-margin business, a twenty-point measurement error is an annoyance you round away. On thin contribution margin, that same error is the line between scaling a winner and quietly funding a loser for a quarter. The honest framing of mmm vs mta vs incrementality isn’t “which one is right” — it’s “which one’s error band fits inside the decision I’m about to make.”
For the neighboring economics, compare Quiz Funnels for Cold Meta Traffic: Zero-Party Payoff and use The Budget Scale-Down Test: DIY Incrementality to validate the measurement decision.
Start with the margin, not the method
The reason measurement error matters at all is the breakeven point it has to clear. Breakeven blended return is roughly 1 ÷ contribution margin. At a 30% contribution margin, you need about a 3.3x return just to wash; at 20%, you need 5x. Those are unforgiving targets, and every measurement method reports a number with a confidence interval around it that you in many cases never see.
Here’s the trap. Say a method tells you a channel is running at 3.6x against a 3.3x breakeven. Looks like a keeper. But if that method’s true error band is plus or minus 20%, the real answer sits somewhere between 2.9x and 4.3x — and the bottom of that range is underwater. The verdict you acted on lived entirely inside the noise. The thinner your margin, the smaller the error you can tolerate, and the smaller the error you can tolerate, the fewer methods qualify for that specific decision.
So the right question for each method is narrow: at the resolution I need, is its error band smaller than the gap between my reported return and my breakeven? Mostly the answer is no, and each method fails that test in a different, predictable place.
MMM: honest at the portfolio level, blurry at the decision in front of you
Marketing mix modeling works top-down. It regresses aggregate outcomes against aggregate spend over time, so it sees everything — paid, organic, promotions, seasonality, price — without touching a single user-level identifier. That’s its real strength, and it’s a durable one as signal loss gets worse.
Where it breaks down is resolution and confidence. MMM speaks in weeks and months, not campaigns and ad sets. When two channels move together — you scale them in lockstep, or you always run them at the same time — the model can’t cleanly separate their effects, and the credit split between them gets unstable. Adstock and carryover assumptions (how long a touch keeps “working”) are modeling choices, and reasonable analysts pick different ones and get different answers. The coefficient on any single channel frequently arrives with an interval wide enough to contain both “scale it” and “cut it.”
MMM breaks down the moment you ask it to adjudicate a tactical move — this ad set, this week, this budget shift. It was never built for that altitude. Trust it to allocate at the portfolio level and to value channels that are hard to track any other way. Don’t let it referee a decision whose entire margin lives in a single campaign.
MTA: precise-looking, structurally biased toward correlation
Multi-touch attribution is the opposite shape: bottom-up, user-level, fast, and granular enough to feel like an answer. It stitches touchpoints into a path and hands out fractional credit. The output looks decision-grade because it’s specific.
Two things rot underneath that specificity. First, the data is increasingly incomplete — consent gating, restricted identifiers, and walled gardens that don’t share paths mean MTA reconstructs journeys from a shrinking, non-random sample of users. The people you can still track aren’t a fair stand-in for everyone. Second, and more fundamental, MTA is correlational by construction. It assigns credit to touches that appeared on the path to a conversion. It cannot tell you which of those conversions would have happened anyway. A retargeting touch served to someone already walking to checkout gets full marks it didn’t earn.
That second flaw is why MTA can overstate exactly the lower-funnel, audience-overlapping tactics that are most straightforward to over-invest in. MTA breaks down whenever the question is causal — “did this spend create demand or just intercept it?” — which, on thin margins, is almost always the question. Use it for path diagnostics, sequencing, and creative-level signal. Don’t read its credit allocation as incremental truth.
Incrementality: the only causal read, and the noisiest one
Holdouts, geo experiments, and lift tests measure the thing the other two can only approximate: outcomes with the spend versus a comparable world without it. That counterfactual is the whole game, and nothing else delivers it.
The catch is statistical power and cost. To detect lift, you withhold spend from a test cell — that withholding has a real opportunity cost while the test runs. And the confidence interval on a lift test is governed by how many conversions land in each cell. When conversions are scarce — small geos, considered purchases, short test windows — the interval gets wide enough that the experiment can’t distinguish a healthy lift from zero. As an illustrative planning range, not a assurance, you in many cases want each test cell to clear a meaningfully high conversion count before the result stabilizes; under that, you’re reading noise with a confidence number stapled to it. Incrementality is also a time-bound snapshot. It measures what you tested, when you tested it, and stops being current the moment creative, audience, or competition shifts.
Incrementality breaks down when conversion volume is too thin to power the test, or when you treat a one-time read as a standing fact. It’s the most trustworthy method and the most straightforward to over-trust.
Where each one stops being trustworthy
| Method | Resolution | Causal? | Breaks down when |
|---|---|---|---|
| MMM | Portfolio, weekly+ | Partly | You ask it to judge a single campaign or short window |
| MTA | Touchpoint, near-real-time | No | The question is “incremental or intercepted?” |
| Incrementality | Tested unit, snapshot | Yes | Conversions per cell are too few to power the test |
The error bands are illustrative, not sourced facts, but the ordering is reliable: MTA looks most precise and is most biased; incrementality is most truthful and most fragile to low volume; MMM is steadiest and least specific.
Triangulate by altitude, not by averaging
Don’t blend the three into one number — averaging a biased estimate, a wide-interval estimate, and a snapshot just launders the error. Assign each to the decision it can actually carry:
- MMM sets the high-level allocation and values the channels you can’t track cleanly.
- Incrementality calibrates the truth — run it on your largest spend lines and use the lift factor to discount the channels MTA flatters.
- MTA handles the fast, tactical loop inside a channel: sequencing, creative, ad-set signal — read directionally, not as causal credit.
The connective tissue is unit economics. Whatever a method reports, push it back through contribution margin and ask whether the answer survives the error band against breakeven. This is the read this kind of tooling should make routine — surfacing the margin-adjusted picture and flagging when a verdict sits inside the noise. Bach AI does exactly that as a read-only layer, and stays read-only until you approve any change.
The takeaway: there’s no single trustworthy method, only methods that are trustworthy at a specific altitude for a specific margin. Before you act on any measurement number, subtract your breakeven and compare what’s left to the method’s error band. If the decision lives inside the noise, you don’t have a measurement problem — you have a stop-and-test problem, and pretending otherwise is how thin-margin accounts bleed.