How to run a geo lift test your budget can read.

A geo lift test changes ad spend in some regions of a country, leaves the others alone and reads the gap in your own revenue. Below: the design for each question, test regions in Germany, France, Italy and the Netherlands, the power maths worked in euros, and what to do when regions can't be split.

By Federico Baravalle, co-founder of Nupact · Platform rules checked October 2026

The short version

A geo lift test (also called a geo holdout or matched market test) changes spend in some regions of one country and compares their revenue with the regions left alone. The question picks the design: go-dark to value spend you already have, holdback for a new channel, heavy-up for the next euro. In Nupact's simulation of a twelve-region country with €60M of annual revenue, an 8-week test on a channel spending €3M a year (5% of that revenue) read about ±1.5 on iROAS, which is directional. One test narrow enough to move budget on, about ±1, took around €4.5M a year (7.5% of revenue). €3M should read about ±1 with a 26-week clean pre-period. At about €1M a year or less (under 2%), every regional design read weak at best: use the platform's lift study or a holdout on your own customer list.

What is a geo lift test?

A geo lift test is a controlled experiment where the unit is a region. You change spend in a set of regions for a fixed number of weeks and compare their sales with a control built from the rest. The gap divided by the spend you changed is the incremental ROAS (iROAS), revenue caused per euro, and because it runs on aggregated sales by region with no user-level tracking, the GDPR footprint stays small.

Attribution cannot stand in for it: against 15 randomised experiments inside Facebook, standard observational methods were off by a factor of three in half the studies (Gordon et al., Marketing Science 2019).

Google and Meta built the two free MMM tools many advertisers start with, and they sell the media being measured. A geo holdout your own team designs is the one number in the stack that its owner has no reason to inflate.

Federico Baravalle, co-founder of Nupact · adapted from a LinkedIn post, August 2026

Which geo test design answers your question?

Google's own geo experiment guidance uses the same three designs (Google Ads Help). At the same spend contrast they are equally precise. They differ in the question and in what the test costs you.

DesignQuestionWhat you give upWhat it measures
Go-darkIs the money we already spend earning its return?Sales in the dark regions. For 8 weeks on a €3M-a-year channel at an average iROAS of about 4: about €370,000 of gross contribution lost at a 40% margin, before the roughly €230,000 of spend savedThe average return of the whole budget; the only design that separates cannibalisation on non-brand search and retargeting
HoldbackDoes this new channel or tactic add anything?A delayed rollout in the held-back regionsThe lift the launch created
Heavy-up +100%Does the next euro pay back?About €230,000 of extra spend for 8 weeks on a €3M channel, at a lower returnThe return above today's spend, which a scaling decision needs

Costs from Nupact's twelve-region simulation (true marginal iROAS of 2; the average of about 4 is what its 16-week go-dark read), September 2026. Approximate; they scale with channel size and margin.

A long go-dark drifts toward the average return, which sits above the marginal one: in our simulation a 16-week go-dark read close to 4 against a marginal truth of 2. A +50% heavy-up is a weak instrument at any size, about 1.8 to 1.9 times wider than a +100% one.

The money is the smaller cost. A go-dark means stopping spend in the test regions for weeks, which is a conversation with the channel team and finance, sometimes with the board. Underfund the test and it returns noise, which people then read as a verdict on the channel.

Federico Baravalle, co-founder of Nupact · adapted from a LinkedIn post, June 2026

How do you choose test regions in European countries?

European countries offer far fewer units than the 210 US media markets, and one capital region often dominates. Use the finest unit both platforms and your sales data support, then group units into 10 to 16 test cells. Outside the US, Google's geo experiments take city IDs or postal codes in the Google Ads interface, and city IDs only through the API (Google Ads Help).

CountryUnitsCells that tend to workWatch out for
Germany16 Bundesländer; 8 Nielsen areas in practice (IIIa and IIIb counted separately); about 400 Kreise12 cells: fold Hamburg into Schleswig-Holstein, Bremen into Niedersachsen, Berlin into Brandenburg, Saarland into Rheinland-Pfalz. Stratify by Nielsen areaEight Nielsen areas are too few to randomise on
France13 metropolitan regions; 96 départementsDépartement clusters stratified by regionNothing matches Île-de-France: keep it in control or leave it out
Italy20 regions; about a hundred provincesProvince clusters stratified north, centre, southA strong north-south gap in conversion rates
Netherlands12 provinces; 40 COROP regionsCOROP regions or four-digit postcode clustersRandstad commuting leaks exposure between cells

Small countries such as Austria, Belgium, Denmark or Portugal are the hard case. Below roughly 15 to 20 usable units there is no honest design on administrative regions: randomise postcode clusters and accept some leakage, which pulls the estimate toward zero, or use a user-level option.

Matched markets or synthetic control: which should you use?

Matched markets pick control regions that tracked the test regions before the test; Google's Time-Based Regression projects that relationship forward and was built for few regions (Kerman, Wang and Vaver). Synthetic control blends untreated regions into a control that reproduces the test regions' history (Abadie 2021); Meta's GeoLift uses it and recommends 20 or more geo units with at least 25 pre-treatment periods. Google's open-source Meridian GeoX designs tests by stratified random sampling and turns results into priors for its MMM (Meridian GeoX).

With twelve cells, randomise inside strata, analyse with Time-Based Regression and use synthetic control as a second opinion. Picking regions by hand invites a choice that already correlates with sales. Report intervals alongside any p-value. Placebo inference is coarse with few regions: with one treated region among 12, the smallest p-value a placebo test can return is 1/12, about 0.083 (Lei and Sudijono 2024); treating 6 of 12, as in the design below, gives 924 possible assignments and a much finer grid. Tools also trade misses for false alarms at this scale: with one treated geo, 20 controls and a true 7.5% lift, GeoLift missed the effect 91.3% of the time and CausalImpact raised false positives 27.8% of the time (Recast simulation).

How long should a geo lift test run?

Eight weeks is a sound default for a mid-market channel. Go to 16 only when the clean pre-period behind it is long, since the pre-period should be three to five times the test, and add a cool-down if conversions lag. Google requires pre-test data of at least three times the test and wants a heavy-up's 4 to 5 day learning period inside the test (Google Ads Help). GeoLift asks for at least 15 days on daily data or 4 to 6 weeks on weekly data, covering one purchase cycle, with 4 to 5 times the test duration of stable history before it (GeoLift best practices).

Length buys less than people expect. In Nupact's simulation, going from 4 to 16 weeks tightened the interval by only 15 to 20%, because twelve regions are the sample and regional shocks that last the whole test do not average out. A longer pre-period did more: on a €1M channel, 26 weeks of pre-period instead of 8 narrowed a 16-week test from ±4.0 to ±2.6, a cut of about a third. If the same cut holds for an 8-week test on a €3M channel, its ±1.5 comes down to about ±1.

How much budget does a geo lift test need?

Channel spend as a share of the revenue you measure decides precision. Nupact simulated a twelve-region country: €60M revenue a year, the top three regions holding half, 6% residual weekly noise per region, an 8-week pre-period, six regions treated, 300 runs. The 95% half-widths on iROAS:

Design, 6 of 12 regions, 4 / 8 / 16 weeks€1M a year€2M a year€3M a year
Go-dark, holdback or +100% heavy-up±4.8 / 4.4 / 4.0±2.4 / 2.2 / 2.0±1.6 / 1.5 / 1.3
Heavy-up +50%±8.8 / 8.1 / 7.1±4.4 / 4.0 / 3.6±2.9 / 2.7 / 2.4

Nupact simulation, September 2026 (heavy-up +100% came within 0.1 of go-dark). Noise matters as much as spend: the €1M four-week go-dark was ±3.3 at 4% noise and ±7.8 at 10%. Full results by design, budget and length, dither included, are in how precise is a geo test.

We call a read decision-grade at about ±1. At that width, one test places a channel on the right side of its break-even with 80% power whenever its true iROAS sits about 1.4 or more away from the line, which covers most channels worth arguing about. On this €60M market, €3M a year is 5% of revenue and reads ±1.5 at 8 weeks, which is directional. The same design reaches about ±1 at around €4.5M a year (7.5% of revenue). €3M would reach about ±1 with a 26-week clean pre-period (an estimate from the simulated €1M run).

The euro figures in this section belong to the €60M reference market. What carries over to yours is the share of revenue: 2%, 5% and 7.5%. That holds only while your regions are about as quiet as the simulation's 6% weekly noise, and in small accounts regional noise usually runs far higher.

Doubling spend halves the interval. Our rule of thumb from the simulation, to use as a first screen before a power analysis on your own pre-period:

Half-width ≈ K × (€1M ÷ annual channel spend) × (annual revenue ÷ €60M) × (noise ÷ 6%), with K = 4.8, 4.4 or 4.0 for a 4, 8 or 16-week go-dark. The minimum detectable effect (MDE), the smallest true iROAS the test tells from zero with 80% power at the 5% level, is about 1.4 times the half-width. The same 1.4 half-widths apply against break-even: to settle whether a channel pays with 80% power, its true iROAS has to sit that far from break-even.

A worked example for Meta prospecting in Germany

Say a retailer with a 40% contribution margin, so a break-even iROAS of 2.5, spends €150,000 a month on Meta prospecting in Germany, in a manual campaign with regional ad sets, on €3,000,000 a month of German online revenue. The plan: an 8-week go-dark in 6 of the 12 German cells, holding about half the revenue.

  1. Annualise. Spend €150,000 × 12 = €1.8M; revenue €3,000,000 × 12 = €36M.
  2. Screen. 4.4 × (1 ÷ 1.8) × (36 ÷ 60) × (6 ÷ 6) = 1.47, so about ±1.5. MDE: 1.4 × 1.5 = 2.1.
  3. Translate into a dip. Spend is 5% of revenue, so a true iROAS of 2 shows as a 10% revenue dip in the dark cells. In euros: €150,000 × 12 ÷ 52 × 8 × 0.5 = €138,000 of spend removed, against €3,000,000 × 12 ÷ 52 × 8 × 0.5 = €2.77M of normal revenue there. At iROAS 2 the gap is €277,000, 10% of €2.77M.
  4. Decide. Against zero, an expected iROAS near 2 sits just under the 2.1 MDE, so the test detects it less than 80% of the time. Against the 2.5 break-even, a ±1.5 interval settles the budget question with 80% power only if the true iROAS is below about 0.4 or above about 4.6 (2.5 ± 2.1). The interval would narrow to about ±1 with a 26-week clean pre-period. Pooling a second test with this one helps further.
Geo test screen · your numbers
95% interval on iROAS
±1.5
Minimum detectable iROAS
2.1
Revenue dip at iROAS 2
−10%

Directional. Plan a longer clean pre-period or a repeat test before moving budget on it.

Screening estimate scaled by spend share from Nupact's twelve-region simulation. It assumes clean regional history and cells that can be switched cleanly. Verdicts: decision-grade up to ±1, directional to ±2.5, weak to ±4, too small for a regional test above that. Under about €80,000 a month of spend or €1M a month of revenue, the screen shows a caution instead, because regional noise at that size usually runs far above the simulation's 6%. A full power analysis runs on your own pre-period.

How do you read a geo lift result?

Back to Germany. After 8 weeks, revenue in the dark cells came in €249,000 below the synthetic control. Spend removed was €138,000, so iROAS = 249,000 ÷ 138,000 = 1.8, and with the ±1.5 interval the read is 1.8, between 0.3 and 3.3. Meta reported a 3.6 ROAS for the same weeks, so its attribution credits the campaign with twice the measured return, an incrementality factor of 0.5 (1.8 ÷ 3.6).

Hold the interval against zero and against break-even (1 ÷ contribution margin, so 2.5 at a 40% margin):

Inconclusive is a normal result at mid-market scale. The interval still bounds the effect, 0.3 to 3.3 in the example, and the next test on the channel starts from that range.

Which mistakes break geo lift tests?

What if your regions can't be split?

Meta Advantage+ sales campaigns would need restructuring into a manual campaign for a regional test, Performance Max leaks across regions, and a channel at about €1M a year or less on €60M (under 2%) reads weak at best anyway. User-level options remain:

OptionRequirements (official)Blind spot
Google Conversion Lift, user-based$5,000 campaign budget and 1,000 observed conversions; 7 days minimum, more than 14 recommended; 1% to 50% holdback; via your rep (Google Ads Help)Other channels; conversions Google does not observe
Meta Conversion LiftRandomised test and control run by Meta; access "is currently limited", and no spend floor is published in Meta's developer docs (Meta for Developers)Halo on search and other channels
TikTok Conversion Lift StudyManaged by account reps; minimum spend depends on objective and scope (TikTok Business Help)Effects outside TikTok
Own-audience holdoutSplit your customer list or site visitors at random, suppress ads to one half on every platform for 2 to 6 weeksProspecting; it reads retargeting and CRM spend

Platform requirements checked October 2026.

Platform studies read at spend levels regions never reach, but the platform measures its own channel on the conversions it observes, so reconcile each read against your own revenue. The own-audience holdout is the cleanest design you control end to end and works on Advantage+ and Performance Max when the exclusion runs on every platform. How these reads combine with a model is covered in incrementality testing vs MMM.

Where Nupact fits

Know what can be measured before you switch anything off.

Nupact does the maths on your own regions, history and spend before anything is promised. Each channel gets one of three verdicts: we can test it, the platform runs it and we check, or too small to test.

In the feasibility table on our incrementality page, Google Search-DE non-brand (31% of spend) gets an eight-week regional go-dark in half the regions with brand search excluded, while TikTok Spark-DE (5%) is too small and falls back to TikTok's own lift study where available, or a benchmark flagged as unmeasured. Bring the regional history of one channel and we run the check for free.

What do people ask before their first geo lift test?

What is the difference between a geo lift test and a geo holdout test?

A geo holdout removes or withholds ads in some regions: a go-dark on an existing channel or a holdback on a new one. "Geo lift" also covers heavy-ups, where spend goes up. Matched market testing is the same family, named after how the control regions are chosen.

How many regions do I need for a geo lift test?

GeoLift recommends 20 or more geo units with at least 25 pre-treatment periods; Time-Based Regression was designed for fewer, and Nupact's simulation used twelve. Below about 15 to 20 usable units, use postcode clusters or a user-level lift study.

Can I run a geo experiment in Google Ads?

Yes. Google Ads geo experiments support holdback, go-dark and heavy-up, on city IDs or postal codes outside the US. Geo-based Conversion Lift goes through your Google rep, needs campaigns that target a single country and runs holdback and go-dark studies only in its beta (Google Ads Help). The open-source Meridian GeoX runs on your own data for any publisher.

Can I geo test Meta Advantage+ sales campaigns?

Only by duplicating the channel into a manual campaign with regional ad sets for the duration of the test. Meta Conversion Lift, or a customer-list holdout excluded from the campaign, is usually cleaner.

What does an inconclusive geo test result mean?

The interval was too wide to place the channel on one side of break-even, though it still bounds the effect. Hold the budget, lengthen the clean pre-period, or plan a repeat to pool with this one.

How much spend does a geo lift test need?

In Nupact's twelve-region simulation, an 8-week go-dark read about ±1.5 on iROAS for a channel spending 5% of the market's revenue (€3M a year on €60M), and about ±1 at 7.5% (€4.5M). €3M would reach about ±1 with a 26-week clean pre-period (an estimate from the simulated €1M run). At about €1M a year or less on €60M (under 2%), no regional design read better than weak: ±4.0 over 16 weeks, or ±2.6 with a 26-week clean pre-period. Your own regional noise moves these lines, so run the screen above on your numbers.