What is a geo lift test?
A geo lift test is a controlled experiment where the unit is a region. You change spend in a set of regions for a fixed number of weeks and compare their sales with a control built from the rest. The gap divided by the spend you changed is the incremental ROAS (iROAS), revenue caused per euro, and because it runs on aggregated sales by region with no user-level tracking, the GDPR footprint stays small.
Attribution cannot stand in for it: against 15 randomised experiments inside Facebook, standard observational methods were off by a factor of three in half the studies (Gordon et al., Marketing Science 2019).
Google and Meta built the two free MMM tools many advertisers start with, and they sell the media being measured. A geo holdout your own team designs is the one number in the stack that its owner has no reason to inflate.
Which geo test design answers your question?
Google's own geo experiment guidance uses the same three designs (Google Ads Help). At the same spend contrast they are equally precise. They differ in the question and in what the test costs you.
| Design | Question | What you give up | What it measures |
|---|---|---|---|
| Go-dark | Is the money we already spend earning its return? | Sales in the dark regions. For 8 weeks on a €3M-a-year channel at an average iROAS of about 4: about €370,000 of gross contribution lost at a 40% margin, before the roughly €230,000 of spend saved | The average return of the whole budget; the only design that separates cannibalisation on non-brand search and retargeting |
| Holdback | Does this new channel or tactic add anything? | A delayed rollout in the held-back regions | The lift the launch created |
| Heavy-up +100% | Does the next euro pay back? | About €230,000 of extra spend for 8 weeks on a €3M channel, at a lower return | The return above today's spend, which a scaling decision needs |
Costs from Nupact's twelve-region simulation (true marginal iROAS of 2; the average of about 4 is what its 16-week go-dark read), September 2026. Approximate; they scale with channel size and margin.
A long go-dark drifts toward the average return, which sits above the marginal one: in our simulation a 16-week go-dark read close to 4 against a marginal truth of 2. A +50% heavy-up is a weak instrument at any size, about 1.8 to 1.9 times wider than a +100% one.
The money is the smaller cost. A go-dark means stopping spend in the test regions for weeks, which is a conversation with the channel team and finance, sometimes with the board. Underfund the test and it returns noise, which people then read as a verdict on the channel.
How do you choose test regions in European countries?
European countries offer far fewer units than the 210 US media markets, and one capital region often dominates. Use the finest unit both platforms and your sales data support, then group units into 10 to 16 test cells. Outside the US, Google's geo experiments take city IDs or postal codes in the Google Ads interface, and city IDs only through the API (Google Ads Help).
| Country | Units | Cells that tend to work | Watch out for |
|---|---|---|---|
| Germany | 16 Bundesländer; 8 Nielsen areas in practice (IIIa and IIIb counted separately); about 400 Kreise | 12 cells: fold Hamburg into Schleswig-Holstein, Bremen into Niedersachsen, Berlin into Brandenburg, Saarland into Rheinland-Pfalz. Stratify by Nielsen area | Eight Nielsen areas are too few to randomise on |
| France | 13 metropolitan regions; 96 départements | Département clusters stratified by region | Nothing matches Île-de-France: keep it in control or leave it out |
| Italy | 20 regions; about a hundred provinces | Province clusters stratified north, centre, south | A strong north-south gap in conversion rates |
| Netherlands | 12 provinces; 40 COROP regions | COROP regions or four-digit postcode clusters | Randstad commuting leaks exposure between cells |
Small countries such as Austria, Belgium, Denmark or Portugal are the hard case. Below roughly 15 to 20 usable units there is no honest design on administrative regions: randomise postcode clusters and accept some leakage, which pulls the estimate toward zero, or use a user-level option.
Matched markets or synthetic control: which should you use?
Matched markets pick control regions that tracked the test regions before the test; Google's Time-Based Regression projects that relationship forward and was built for few regions (Kerman, Wang and Vaver). Synthetic control blends untreated regions into a control that reproduces the test regions' history (Abadie 2021); Meta's GeoLift uses it and recommends 20 or more geo units with at least 25 pre-treatment periods. Google's open-source Meridian GeoX designs tests by stratified random sampling and turns results into priors for its MMM (Meridian GeoX).
With twelve cells, randomise inside strata, analyse with Time-Based Regression and use synthetic control as a second opinion. Picking regions by hand invites a choice that already correlates with sales. Report intervals alongside any p-value. Placebo inference is coarse with few regions: with one treated region among 12, the smallest p-value a placebo test can return is 1/12, about 0.083 (Lei and Sudijono 2024); treating 6 of 12, as in the design below, gives 924 possible assignments and a much finer grid. Tools also trade misses for false alarms at this scale: with one treated geo, 20 controls and a true 7.5% lift, GeoLift missed the effect 91.3% of the time and CausalImpact raised false positives 27.8% of the time (Recast simulation).
How long should a geo lift test run?
Eight weeks is a sound default for a mid-market channel. Go to 16 only when the clean pre-period behind it is long, since the pre-period should be three to five times the test, and add a cool-down if conversions lag. Google requires pre-test data of at least three times the test and wants a heavy-up's 4 to 5 day learning period inside the test (Google Ads Help). GeoLift asks for at least 15 days on daily data or 4 to 6 weeks on weekly data, covering one purchase cycle, with 4 to 5 times the test duration of stable history before it (GeoLift best practices).
Length buys less than people expect. In Nupact's simulation, going from 4 to 16 weeks tightened the interval by only 15 to 20%, because twelve regions are the sample and regional shocks that last the whole test do not average out. A longer pre-period did more: on a €1M channel, 26 weeks of pre-period instead of 8 narrowed a 16-week test from ±4.0 to ±2.6, a cut of about a third. If the same cut holds for an 8-week test on a €3M channel, its ±1.5 comes down to about ±1.
How much budget does a geo lift test need?
Channel spend as a share of the revenue you measure decides precision. Nupact simulated a twelve-region country: €60M revenue a year, the top three regions holding half, 6% residual weekly noise per region, an 8-week pre-period, six regions treated, 300 runs. The 95% half-widths on iROAS:
| Design, 6 of 12 regions, 4 / 8 / 16 weeks | €1M a year | €2M a year | €3M a year |
|---|---|---|---|
| Go-dark, holdback or +100% heavy-up | ±4.8 / 4.4 / 4.0 | ±2.4 / 2.2 / 2.0 | ±1.6 / 1.5 / 1.3 |
| Heavy-up +50% | ±8.8 / 8.1 / 7.1 | ±4.4 / 4.0 / 3.6 | ±2.9 / 2.7 / 2.4 |
Nupact simulation, September 2026 (heavy-up +100% came within 0.1 of go-dark). Noise matters as much as spend: the €1M four-week go-dark was ±3.3 at 4% noise and ±7.8 at 10%. Full results by design, budget and length, dither included, are in how precise is a geo test.
We call a read decision-grade at about ±1. At that width, one test places a channel on the right side of its break-even with 80% power whenever its true iROAS sits about 1.4 or more away from the line, which covers most channels worth arguing about. On this €60M market, €3M a year is 5% of revenue and reads ±1.5 at 8 weeks, which is directional. The same design reaches about ±1 at around €4.5M a year (7.5% of revenue). €3M would reach about ±1 with a 26-week clean pre-period (an estimate from the simulated €1M run).
The euro figures in this section belong to the €60M reference market. What carries over to yours is the share of revenue: 2%, 5% and 7.5%. That holds only while your regions are about as quiet as the simulation's 6% weekly noise, and in small accounts regional noise usually runs far higher.
Doubling spend halves the interval. Our rule of thumb from the simulation, to use as a first screen before a power analysis on your own pre-period:
Half-width ≈ K × (€1M ÷ annual channel spend) × (annual revenue ÷ €60M) × (noise ÷ 6%), with K = 4.8, 4.4 or 4.0 for a 4, 8 or 16-week go-dark. The minimum detectable effect (MDE), the smallest true iROAS the test tells from zero with 80% power at the 5% level, is about 1.4 times the half-width. The same 1.4 half-widths apply against break-even: to settle whether a channel pays with 80% power, its true iROAS has to sit that far from break-even.
A worked example for Meta prospecting in Germany
Say a retailer with a 40% contribution margin, so a break-even iROAS of 2.5, spends €150,000 a month on Meta prospecting in Germany, in a manual campaign with regional ad sets, on €3,000,000 a month of German online revenue. The plan: an 8-week go-dark in 6 of the 12 German cells, holding about half the revenue.
- Annualise. Spend €150,000 × 12 = €1.8M; revenue €3,000,000 × 12 = €36M.
- Screen. 4.4 × (1 ÷ 1.8) × (36 ÷ 60) × (6 ÷ 6) = 1.47, so about ±1.5. MDE: 1.4 × 1.5 = 2.1.
- Translate into a dip. Spend is 5% of revenue, so a true iROAS of 2 shows as a 10% revenue dip in the dark cells. In euros: €150,000 × 12 ÷ 52 × 8 × 0.5 = €138,000 of spend removed, against €3,000,000 × 12 ÷ 52 × 8 × 0.5 = €2.77M of normal revenue there. At iROAS 2 the gap is €277,000, 10% of €2.77M.
- Decide. Against zero, an expected iROAS near 2 sits just under the 2.1 MDE, so the test detects it less than 80% of the time. Against the 2.5 break-even, a ±1.5 interval settles the budget question with 80% power only if the true iROAS is below about 0.4 or above about 4.6 (2.5 ± 2.1). The interval would narrow to about ±1 with a 26-week clean pre-period. Pooling a second test with this one helps further.
Directional. Plan a longer clean pre-period or a repeat test before moving budget on it.
Screening estimate scaled by spend share from Nupact's twelve-region simulation. It assumes clean regional history and cells that can be switched cleanly. Verdicts: decision-grade up to ±1, directional to ±2.5, weak to ±4, too small for a regional test above that. Under about €80,000 a month of spend or €1M a month of revenue, the screen shows a caution instead, because regional noise at that size usually runs far above the simulation's 6%. A full power analysis runs on your own pre-period.
How do you read a geo lift result?
Back to Germany. After 8 weeks, revenue in the dark cells came in €249,000 below the synthetic control. Spend removed was €138,000, so iROAS = 249,000 ÷ 138,000 = 1.8, and with the ±1.5 interval the read is 1.8, between 0.3 and 3.3. Meta reported a 3.6 ROAS for the same weeks, so its attribution credits the campaign with twice the measured return, an incrementality factor of 0.5 (1.8 ÷ 3.6).
Hold the interval against zero and against break-even (1 ÷ contribution margin, so 2.5 at a 40% margin):
- Entirely above break-even: the channel pays; a heavy-up is the next test.
- Below break-even, above zero: it causes sales but loses money at this spend; cut toward the level where the next euro pays.
- Straddling break-even: our example, where Meta prospecting clearly causes sales and the test cannot say whether it pays at 2.5. Hold the budget and plan a repeat that can be pooled with this one.
- Including zero: the test could not tell the channel from nothing. Before cutting, check whether its MDE was ever below the effect you expected.
Inconclusive is a normal result at mid-market scale. The interval still bounds the effect, 0.3 to 3.3 in the example, and the next test on the channel starts from that range.
Which mistakes break geo lift tests?
- Letting the platform re-spend the cut. Under an Advantage+ campaign budget Meta reallocates between ad sets itself. On Google, Performance Max picks up queries a paused Search campaign leaves. Search takes priority over Performance Max only when one of its keywords, of any match type, is identical or spell-corrected to the query, and Ad Rank decides the rest (Google Ads Help).
- Measuring platform conversions. Read first-party revenue or orders by region; the platform cannot see the regions where it stopped serving.
- Stopping early because the line looks good, which inflates false positives.
- Testing through Black Friday or a regional holiday, when a shock that hits some cells harder than others gets read as an ad effect.
- Leaving Google's default "Presence or interest" location setting on. Set "Presence" where the campaign type allows it (Google Ads Help).
- Deciding the analysis after seeing the data. Write down the outcome, length, cool-down, method and break-even line before launch, then read once at the planned end.
- Promising a result before the power check, then calling the inconclusive read a zero.
What if your regions can't be split?
Meta Advantage+ sales campaigns would need restructuring into a manual campaign for a regional test, Performance Max leaks across regions, and a channel at about €1M a year or less on €60M (under 2%) reads weak at best anyway. User-level options remain:
| Option | Requirements (official) | Blind spot |
|---|---|---|
| Google Conversion Lift, user-based | $5,000 campaign budget and 1,000 observed conversions; 7 days minimum, more than 14 recommended; 1% to 50% holdback; via your rep (Google Ads Help) | Other channels; conversions Google does not observe |
| Meta Conversion Lift | Randomised test and control run by Meta; access "is currently limited", and no spend floor is published in Meta's developer docs (Meta for Developers) | Halo on search and other channels |
| TikTok Conversion Lift Study | Managed by account reps; minimum spend depends on objective and scope (TikTok Business Help) | Effects outside TikTok |
| Own-audience holdout | Split your customer list or site visitors at random, suppress ads to one half on every platform for 2 to 6 weeks | Prospecting; it reads retargeting and CRM spend |
Platform requirements checked October 2026.
Platform studies read at spend levels regions never reach, but the platform measures its own channel on the conversions it observes, so reconcile each read against your own revenue. The own-audience holdout is the cleanest design you control end to end and works on Advantage+ and Performance Max when the exclusion runs on every platform. How these reads combine with a model is covered in incrementality testing vs MMM.
Know what can be measured before you switch anything off.
Nupact does the maths on your own regions, history and spend before anything is promised. Each channel gets one of three verdicts: we can test it, the platform runs it and we check, or too small to test.
In the feasibility table on our incrementality page, Google Search-DE non-brand (31% of spend) gets an eight-week regional go-dark in half the regions with brand search excluded, while TikTok Spark-DE (5%) is too small and falls back to TikTok's own lift study where available, or a benchmark flagged as unmeasured. Bring the regional history of one channel and we run the check for free.
What do people ask before their first geo lift test?
What is the difference between a geo lift test and a geo holdout test?
A geo holdout removes or withholds ads in some regions: a go-dark on an existing channel or a holdback on a new one. "Geo lift" also covers heavy-ups, where spend goes up. Matched market testing is the same family, named after how the control regions are chosen.
How many regions do I need for a geo lift test?
GeoLift recommends 20 or more geo units with at least 25 pre-treatment periods; Time-Based Regression was designed for fewer, and Nupact's simulation used twelve. Below about 15 to 20 usable units, use postcode clusters or a user-level lift study.
Can I run a geo experiment in Google Ads?
Yes. Google Ads geo experiments support holdback, go-dark and heavy-up, on city IDs or postal codes outside the US. Geo-based Conversion Lift goes through your Google rep, needs campaigns that target a single country and runs holdback and go-dark studies only in its beta (Google Ads Help). The open-source Meridian GeoX runs on your own data for any publisher.
Can I geo test Meta Advantage+ sales campaigns?
Only by duplicating the channel into a manual campaign with regional ad sets for the duration of the test. Meta Conversion Lift, or a customer-list holdout excluded from the campaign, is usually cleaner.
What does an inconclusive geo test result mean?
The interval was too wide to place the channel on one side of break-even, though it still bounds the effect. Hold the budget, lengthen the clean pre-period, or plan a repeat to pool with this one.
How much spend does a geo lift test need?
In Nupact's twelve-region simulation, an 8-week go-dark read about ±1.5 on iROAS for a channel spending 5% of the market's revenue (€3M a year on €60M), and about ±1 at 7.5% (€4.5M). €3M would reach about ±1 with a 26-week clean pre-period (an estimate from the simulated €1M run). At about €1M a year or less on €60M (under 2%), no regional design read better than weak: ±4.0 over 16 weeks, or ±2.6 with a 26-week clean pre-period. Your own regional noise moves these lines, so run the screen above on your numbers.