Why does attribution disagree with incrementality tests?
Attribution shares out credit for conversions that happened. Incrementality asks how many of those conversions would have happened anyway without the ad, and the gap between the two answers is largest where ads reach people already on their way to buy. When eBay switched off brand search ads, 99.5% of the paid clicks came back through organic results (Blake, Nosko and Tadelis, Econometrica 2015).
Every multi-touch model, Markov chains and Shapley values included, divides credit for conversions that happened. Buyers already on their way to converting collect the most late-funnel touches, so path-based credit flatters whatever sits closest to the purchase. Branded search has collected that credit for two decades.
Across 663 Facebook experiments, even a model using double machine learning on more than 5,000 user features put the median lift for upper-funnel outcomes at 83%, where the randomised tests measured 29% (Gordon, Moakler and Zettelmeyer 2023).
Keep attribution for speed and granularity, and scale it by a measured factor: incrementality factor = tested iROAS ÷ attributed ROAS for the same channel and period. From our product demo: Meta Retargeting-DE shows 7,600 conversions in 30 days under your attribution, and an audience holdout put its factor at 0.25, so 7,600 × 0.25 = 1,900 incremental.
Conversion lift vs geo lift: which should you use?
A conversion lift study randomises users inside one platform, so its sample is millions of people and it reads at spend levels where regions are hopeless. Google's user-based study needs a $5,000 campaign budget and 1,000 observed conversions, runs at least 7 days with more than 14 recommended, and treats results at 50% to 90% study power as directional (Google Ads Help). Access to Meta's Conversion Lift "is currently limited" (Meta for Developers), and the docs point you to a Meta representative for access. TikTok runs its study as a managed service with minimums set per objective (TikTok Business Help).
A geo lift test changes spend by region and reads your own revenue, so it sees what the platform cannot, such as halo on search and cannibalisation between channels. The price is that you need regions you can split cleanly and enough spend in each market. The geo lift test guide works through the sizing in euros.
- Pick conversion lift for Meta Advantage+ sales and Google Performance Max, for channels with a small share of a market's revenue (about €1M a year or less on €60M, under 2%), and for small countries with too few regions.
- Pick a geo test when the decision is large and cross-channel, when you need the answer in first-party revenue, or when you doubt a non-brand search or retargeting line that may be cannibalising.
- Run both on the same channel once a year where you can. When they disagree, the geo read on your own revenue wins, and the gap tells you how much to discount the platform's next study.
What can a marketing mix model tell you at mid-market scale?
An MMM is the only method that covers every channel, including the small ones, and draws saturation curves you can plan a budget on. Its weakness is identification: Google's own researchers found five models fitted to the same data, all with R-squared of 0.98 to 0.99, that differed in sales predictions by up to 50% (Chan and Perry, Google 2017). Meridian's own documentation works through a national model with 12 channels and 26 parameters: two years of weekly data gives about four data points per parameter, which it calls "too low to estimate the model reliably", and it points to regional data as the fix (Meridian docs).
A typical European mid-market advertiser has 12 to 24 months of clean history and six or more correlated channels. Uncalibrated, the model will produce a number for each of them, with an interval wide enough to justify almost any budget. Keep the tested iROAS as the headline return and use the model as the forecast between tests.
How do you calibrate an MMM with experiments?
Every major open-source MMM takes experiment results as input. Meridian turns them into ROI priors, with a CalibrationBuilder that automates the translation, and warns that "experiments and MMMs often have different estimands" (Meridian docs). Robyn adds the gap between model and experiment as an extra objective, cites a third-party finding that uncalibrated models show a 25% average difference to ground truth, and a simulation in which accuracy improved with up to 10 studies per channel (Robyn docs). PyMC-Marketing adds each lift test as a likelihood term on the channel's saturation curve (PyMC-Marketing). In a Google simulation under idealised conditions, calibration made the posterior on ROAS about five times tighter (Zhang et al., Google 2024).
- Test the channels that carry the budget first. Two or three large channels per market, read by geo test or platform lift.
- Enter each read with its interval. Widen platform lift reads before they go in, because the platform measured its own channel on its own conversions.
- Match the estimand. A go-dark reads the average return, a heavy-up the return above today's spend, and the MMM's ROI is against zero spend. Say which one each prior is.
- Let old tests weigh less. A test from last spring describes last spring's channel, so its pull on the model should fade as it ages.
- Retest where the model and the last test part ways, or where the test has aged past a season.
Which method fits your spend level?
Spend per channel per market decides what can be measured, since precision follows the channel's spend as a share of the revenue in that market. These tiers come from Nupact's precision analysis of twelve-region European countries (September 2026) and are approximate.
| Your situation | Run | Skip |
|---|---|---|
| A channel spending about €1M a year or less in one market (under about 2% of that market's revenue in our simulation) | That platform's lift study; a customer-list holdout if it is retargeting; a benchmark clearly labelled unmeasured | Regional tests: none reads better than weak |
| €2M to €10M digital a year in one country, most channels under €2M | Platform lift studies on Meta and Google about once a quarter, customer-list holdouts, a light MMM with experiment priors, and at most one regional test a year on the biggest channel, treated as directional | An uncalibrated MMM as the budget number |
| €10M or more in one country, two or three channels at €3M or more (5% or more of that market's revenue) | Three to five regional tests a year of 8 to 16 weeks, rotating across the readable channels, each on a 26-week clean pre-period, which should take an 8-week test at €3M from about ±1.5 to about ±1 (an estimate from the simulated €1M run); a calibrated MMM refreshed weekly; attribution scaled by each channel's latest factor | Testing every small channel |
| Several European markets | Qualify market by market. A large market's read can serve a small one as a discounted prior until the small market has its own test | Pooling regions across countries without checking that the channel works the same in each |
How do you make incrementality testing always-on?
Every "always-on" incrementality number on the market is a periodic test with a model or a scaled report filling the weeks between reads. Build the system around that, with a test's result applied daily and its age shown next to the number.
- Daily: each channel's attributed numbers multiplied by its factor from the latest test, labelled with the test and its age. As the test ages, the factor moves toward the model's estimate.
- Weekly: the calibrated MMM's forecast, with an interval that widens as the last test gets older.
- Quarterly: platform lift studies on Meta and Google as cheap calibration points, reconciled against first-party revenue.
- A few times a year: regional tests on the channels big enough to read, planned in advance on a calendar so the large channels take turns.
Many tests at mid-market scale come back inconclusive. A standing system records the narrower range as a result and books the next test.
In more than 20 interviews with performance leaders this summer, I could count on one hand the teams that run incrementality seriously enough to move budget with it. The blockers were organisational. Someone in the C-suite has to commit to it as a standing decision. The first honest read always contradicts attribution, so a real person's channel loses its number. And the result lives at portfolio level while budgets move campaign by campaign.
Run incrementality as part of the system that runs your budget.
Nupact runs the regional tests and own-audience holdouts itself, choosing go-dark, holdback or heavy-up by the question. Lift studies are run by the platforms and checked by Nupact against your revenue, and a calibrated model carries the weeks between reads.
The output is a corrected number per channel, every day, with the test it came from and its age. From our product demo: Google Search-DE non-brand shows 9,140 conversions in your attribution over 30 days, a go-dark nine weeks ago measured a factor of 0.75, so the corrected figure is about 6,860, and the budget proposals and alerts after that start from it.
What do people ask when choosing between MMM and lift tests?
Is MMM better than incrementality testing?
Tests measure what one channel caused at one point in time. An MMM estimates every channel continuously but cannot separate correlated channels on 12 to 24 months of data without experimental priors, so at mid-market scale you need tests to calibrate the model and the model to cover the weeks between tests.
What is the difference between conversion lift and geo lift?
Conversion lift randomises users inside one platform and reports that platform's conversions. Geo lift changes spend by region and reads your own revenue, including effects on other channels. Conversion lift reads at much smaller spend than a regional test can, which makes it the default for channels with a small share of a market's revenue (about €1M a year or less on €60M, under 2%).
Can I trust Meta's or Google's conversion lift results?
As a calibration point, yes. The randomisation is sound, but the platform measures its own channel on the conversions it observes and you cannot audit the assignment. Reconcile each read against first-party revenue for the same period, and where a channel has both a lift study and a geo test, trust the geo test.
How often should you run incrementality tests?
At mid-market scale, three to five regional tests a year on the two or three channels big enough to read, plus a platform lift study per major platform about every quarter. Retest any channel whose test is more than a season old or whose spend has changed a lot.
What is the difference between incrementality and attribution?
Attribution assigns credit for conversions that happened to the ads people touched. Incrementality measures how many conversions the ads caused, by comparing a randomised group that could see the ads with one that could not. Retargeting and brand search usually look far better in attribution than in a test.
Should I use Meridian or Robyn?
Both are free and take experiment results. Meridian is Bayesian, models regions natively and pairs with Google's open-source GeoX for geo tests. Robyn uses ridge regression with an evolutionary search and calibrates against lift studies through an extra objective. With few years of data, the quality of the experiments you feed in matters more than the choice of library.