How wide is the interval for each design, budget and length?
Each figure is the 95% half-width on iROAS, the incremental revenue per euro of spend, in the same units. A read of 2 at ±4.4 puts the true value anywhere from −2.4 to 6.4, so the test cannot say whether the channel earns anything.
| Test length | €1M a year 1.7% of revenue | €2M a year 3.3% of revenue | €3M a year 5% of revenue |
|---|---|---|---|
| Go-dark or holdback · 6 of 12 regions | |||
| 4 weeks | ±4.8 | ±2.4 | ±1.6 |
| 8 weeks | ±4.4 | ±2.2 | ±1.5 |
| 16 weeks | ±4.0 | ±2.0 | ±1.3 |
| Heavy-up +100% · 6 of 12 regions | |||
| 4 weeks | ±4.7 | ±2.3 | ±1.6 |
| 8 weeks | ±4.3 | ±2.2 | ±1.4 |
| 16 weeks | ±3.9 | ±1.9 | ±1.3 |
| Heavy-up +50% · 6 of 12 regions | |||
| 4 weeks | ±8.8 | ±4.4 | ±2.9 |
| 8 weeks | ±8.1 | ±4.0 | ±2.7 |
| 16 weeks | ±7.1 | ±3.6 | ±2.4 |
| Dither ±20%, weekly sign changes · all 12 regions | |||
| 4 weeks | ±25.1 | ±12.5 | ±8.4 |
| 8 weeks | ±15.0 | ±7.5 | ±5.0 |
| 16 weeks | ±9.3 | ±4.6 | ±3.1 |
| Dither ±20%, 4-week sign holds · all 12 regions | |||
| 4 weeks | n/a | n/a | n/a |
| 8 weeks | ±29.6 | ±14.8 | ±9.8 |
| 16 weeks | ±9.3 | ±4.7 | ±3.1 |
Nupact simulation, September 2026. Go-dark and holdback share a row because both compare regions with and without the channel. Dither moves budget up 20% in one region of each pair and down in the other, so national spend stays flat. With 4-week holds, a 4-week test is a single pulse and gives no usable estimate.
No cell reaches ±1, the width we call decision-grade. At ±1, one test places a channel on the right side of break-even with 80% power when its true iROAS sits about 1.4 or more from the line. The closest cell is €3M a year over 16 weeks, at ±1.3.
On your own market, read each column by its share of revenue, since that share is what carries over. The interval also grows with your regional noise. The geo test calculator does the scaling for you.
What do the results mean for your test plan?
Go-dark and a +100% heavy-up are equally precise
Both create the same spend gap between test and control regions, and their intervals sat within about 0.1 of each other in every cell. A +50% heavy-up halves that gap and came out about 1.8 times as wide. Which of the first two to run depends on the question, as the geo lift guide sets out.
Small budget pulses are too noisy to read within 16 weeks
A ±20% weekly dither read ±25.1 on a €1M channel over 4 weeks and still ±9.3 over 16. At €3M a year and 16 weeks it was ±3.1, more than twice the go-dark's ±1.3. In an earlier run of the same model, a weekly ±20% dither on a €1.2M channel needed about a year to match one 4-week go-dark.
Add pre-period weeks before you add test weeks
For the €1M go-dark, lengthening the clean pre-period from 8 to 26 weeks narrowed a 4-week test from ±4.8 to ±3.8, and a 16-week test from ±4.0 to ±2.6. Quadrupling the test itself, on 8 weeks of history, only went from ±4.8 to ±4.0. The pre-period is history you already hold, as long as no other test ran in it, while every extra test week costs sales in the treated regions.
How was this simulated, and what does it leave out?
- Market. Weekly revenue in twelve regions with seasonality, a Q4 peak, a 5% annual trend, regional shocks that persist from week to week and random regional promotions. The tested channel has diminishing returns and a carryover half-life of 1.5 weeks.
- Estimator. For the regional designs, a regression of region-week revenue on channel spend with region and week fixed effects, over the pre-period and the test. The dither used an instrumented regression that was told the true carryover, which flatters it.
- The ± figure. The standard deviation of the 300 estimates in each cell times 2.20, the 95% critical value of a t-distribution with 11 degrees of freedom. With 300 runs, each figure carries about 4% simulation error. Cells of the same design family share their random draws, so gaps between them come from the design.
- Bias. The ± measures spread around each design's own average. As a go-dark lengthens, its average drifts from the marginal return of 2 toward the channel's average return of 6.7: +0.9 at 8 weeks, +2.1 at 16. Heavy-ups landed 0.6 to 1.0 below 2, because they read the return on spend above today's level.
- Your noise. The €1M 4-week go-dark read ±3.3 at 4% regional noise and ±7.8 at 10%. Measure your own on your pre-period.
- Spillover. The grid has none. Commuters, IP-based targeting and Performance Max's location setting leak exposure between regions. In a separate run of the same model, 40% leakage between regions pulled the go-dark estimate 70% toward zero.
- Platform friction. Learning-phase resets are not modelled, and Meta Advantage+ sales campaigns cannot be split by region without restructuring the channel. On Google, a dither has to move tROAS or tCPA targets, since budget changes trigger learning.
- Region count. Countries with 30 to 40 usable regions, or quieter regional revenue, read tighter.
Check what each channel can measure before you test it.
The geo lift guide covers design, region choice and reading a result, and its calculator scales these figures to your spend, revenue and noise. For how tests and a model share the work, see incrementality testing vs MMM.
Nupact runs regional tests and holdouts. Lift studies are run by the platforms, and Nupact checks each one against your own revenue. The feasibility check on our incrementality page tells you which applies to each channel.