Geo Incrementality Testing

How to Run a Geo Test on AI Campaigns Without Pausing Spend

How to run geo testing without pausing revenue: use a heavy-up test on Performance Max, Advantage+ or Smart Bidding, size the added spend and read marginal iROAS.

On this page

To run geo testing without pausing revenue, raise spend in the test markets and leave every other market on its current budget. Google's documentation calls this design a heavy-up. For campaigns that Performance Max, Advantage+ or Smart Bidding run, we think it should be the default geo test, with go-dark kept for the narrower cases covered at the end. In our view the decision behind this question is usually whether to scale, and a heavy-up measures the dollars that decision would add.

What a scale-up test measures that a holdout cannot

A "scale-up" geo test instead of a holdout answers a different question from the one a dark market answers. Google Ads Help lists three geo study types. Go-dark exists to "Completely shut off spend in test geos to definitively prove the value of existing, live campaigns." Heavy-up exists to "Inject incremental budget to existing campaigns to forecast the revenue impact of scaling", and the page's example is a Performance Max budget raised in test regions. Holdback covers a tactic you have not launched yet.

In the 2011 geo experiments paper from Google Research, Jon Vaver and Jim Koehler describe the treatment as a spend differential: "This differential will be negative if the campaign change causes the ad spend to decrease in the treatment geos (e.g. campaigns turned off), and positive if the change causes an increase in ad spend (e.g. bids increased or keywords added)." A scale-up test is the same experiment with the sign flipped, and the answer means something different: go-dark estimates the return on the budget you already spend, and heavy-up estimates the incremental return on the next tranche.

How much revenue each design puts at risk

Minimizing revenue loss in geo testing is mostly settled before the holdout begins, by the choice of design. The table uses a hypothetical account: a Performance Max campaign spending $300,000 over a four-week test, test markets that carry 20% of that spend, and an assumed true iROAS of 3.0 on existing spend. That last figure is what a real test would be trying to learn, so read the revenue column as illustration.

DesignSpend change in test marketsRevenue exposed at the assumed 3.0
Go-dark, cell with 20% of spend-$60,000$180,000
Small-cell go-dark, 5% of spend-$15,000$45,000
Partial cut, 50% off in the 20% cell-$30,000Up to $90,000
Heavy-up, +50% in the 20% cell+$30,000$0 of existing revenue; the $30,000 added is the stake

The partial cut's figure is a ceiling, since the dollars cut first are plausibly the weakest. A small cell protects revenue at the price of sensitivity, and only a power analysis says how small is too small. In the simulated data of the GeoLift walkthrough from Meta, the chosen test cell was two cities holding about 3.3% of conversions, with a 10% lift as the minimum for a well-powered 15-day test. That is one simulated dataset, a demonstration of method. Any dark cell carved out of a budget-capped campaign also needs a matching budget reduction: Google's geo experiments page warns that removing test geos otherwise shifts unspent budget into control regions and inflates the baseline.

How much spend to add

Geo testing in advertising without pausing revenue moves the cost of the test from lost sales to added budget, so the sizing question is how much to add. Google's setup example raises the test campaign's daily budget "by 50% relative to its historical run rate", offered as an example.

Precision comes first. Equation 4 of the Vaver and Koehler paper puts the squared deviations of the spend differentials from their mean in the denominator of the variance of the ROAS estimate. By our reading, doubling the differential in every test geo roughly halves the confidence interval. A +50% heavy-up moves half the dollars a go-dark would in the same markets, so it starts with an interval about twice as wide, and the paper names experiment length and test fraction as the other levers.

Budget comes second. GeoLift estimates the investment a well-powered test needs from a cost per incremental conversion, which its walkthrough says can come from marketing mix models, Conversion Lift studies and earlier geo experiments. When the response metric is new customers, that figure is an incremental CAC. Our rule: add enough that the smallest marginal iROAS the test can detect sits below your break-even ROAS.

Setting it up on campaigns the AI runs

Geo-based holdout tests can validate AI campaign performance without pausing revenue, though the built-in tooling stops short of the heavy-up. The Conversion Lift based on geography setup page in Google Ads Help says "You can only run holdback and go-dark studies in this beta", and notes that Conversion Lift is not available for all Google Ads accounts and is requested through a Google account representative. A heavy-up is a manual build in the Google Ads account. Google's open-source Meridian GeoX library, documented on Google for Developers, lists heavy-up among the experiment designs it supports.

Google's geo experiments page says "Location bid modifiers are ignored by Smart Bidding and cross-network formats", which closes the regional bid adjustment route for any campaign on Smart Bidding. The documented route, required for Smart Bidding, cross-network formats and budget-constrained campaigns, is duplication: one copy targets control geos at the historical run rate, the other targets test geos with the raised budget, and both use Presence targeting. For uncapped budgets Google says to loosen the Target CPA or Target ROAS on the test copy, and it rules out heavy-up on uncapped Max Conversions or Maximize Conversion Value bidding without a target. The same page states that a budget change above 20% starts a new learning period, that duplication triggers one as well, and that the typical four to five day learning period must sit inside a heavy-up's test period.

On Meta, we found no help page that walks through a geo heavy-up for Advantage+ campaigns. What we did find is the GeoLift documentation, where Meta describes GeoLift as a research tool for cases where people-based measurement is not feasible and recommends people-based experiments where possible. The Meta Business Help Center says "During the learning phase, ad sets are less stable and usually have a higher CPA", and that ad sets usually exit after about 50 results in the week after the last significant edit. We would add test days to cover it.

Reading marginal iROAS

A geo holdout test of incrementality in marketing, run without pausing campaigns, returns a marginal number, and the benchmark for it is break-even. Google Ads Help defines iROAS as incremental conversion value divided by incremental cost, and incremental cost as the difference in cost between treatment and control during the experiment. In a heavy-up that denominator is the added spend alone.

In the hypothetical account, the test added $30,000 and measured $54,000 of incremental conversion value, a marginal iROAS of 1.8 with an interval of 1.2 to 2.4. At an assumed 40% gross margin, break-even ROAS is 1 divided by 0.40, or 2.5. The whole interval sits under break-even, so the added dollars lose money on first-order margin whatever blended ROAS the platform reports. An interval sitting wholly above 2.5 is the evidence to collect before you scale the budget, and one that straddles 2.5 means the test was too small to decide. Google recommends an optional cooldown window when the conversion cycle runs longer than a few weeks. Our ROAS calculator is a further resource for the break-even arithmetic.

Matched markets still decide whether the result holds

Geo matched market testing keeps the same design rules whether the holdout campaigns go dark or, without pausing revenue, the test markets take a heavier budget. Vaver and Koehler assign geos to treatment and control at random and report narrower intervals when geos are ranked by size before assignment. Google's implementation page asks for pre-test data covering at least three times the test period, no parallel user-level lift studies on the same campaigns, and as few mid-test changes as possible. GeoLift's walkthrough suggests a test period that holds at least one full purchase cycle.

One cost survives every design, the heavy-up included. Google notes that moving from country targeting to city or ZIP code targeting reduces overall traffic, because granular location cannot be determined for every user, and that geos left out of the design receive no business-as-usual budget once campaigns are duplicated.

When a full holdout is still the right test

Go-dark earns its cost in three situations. The first is a decision about keeping or cutting the base budget, because the return on new dollars says little about the first ones, and Google reserves go-dark for proving the value of "existing, live campaigns". The second is a campaign that cannot absorb more spend, such as the uncapped, no-target bidding case above. The third is a heavy-up that came back under break-even, which leaves the base in doubt.

For those cases Google documents a budget-neutral multi-cell design with control, heavy-up and go-dark cells, where spend in the dark geos falls to $0 while the heavy-up geos take an increase. Keep the dark cell as small as the power analysis allows. If you want a second opinion on a design, our paid media team is a further resource. We would still run the heavy-up first: it prices the decision you are about to make, with budget you were already considering.

Sources

Frequently asked questions

Can you run a geo test without turning any campaigns off?
Yes. A heavy-up test raises spend in the test markets and leaves control markets on their current budgets, so no market goes dark. Google Ads Help describes heavy-up as the geo study type for forecasting the revenue impact of scaling. What you put at risk is mainly the added budget, and what you learn is the marginal return on those added dollars.
How much budget should a heavy-up geo test add?
Enough that the smallest marginal iROAS the test can detect sits below your break-even ROAS. Google's setup example uses a 50% increase over the historical run rate, as an illustration only. Google also says a budget change above 20% starts a new learning period, so plan the test length around the typical four to five day learning period.
Does Google Ads Conversion Lift support heavy-up tests?
Google's setup page for Conversion Lift based on geography says "You can only run holdback and go-dark studies in this beta." Conversion Lift is also not available for all Google Ads accounts and is requested through a Google account representative. A heavy-up is built by hand, which for Smart Bidding campaigns means duplicating the campaign into control and test copies.
Is marginal iROAS the same as the ROAS my ad platform reports?
No. Marginal iROAS from a heavy-up divides incremental conversion value by incremental cost, which is the added spend alone. Platform ROAS divides attributed conversion value by all spend. Read the marginal figure against break-even ROAS, which is 1 divided by gross margin. In our hypothetical example, a marginal iROAS of 1.8 falls short of a 2.5 break-even at a 40% margin.
When is a go-dark holdout still necessary?
When the decision is whether the existing budget is worth keeping, when the campaign cannot absorb more spend, or when a heavy-up comes back under break-even and the base is in doubt. Google describes go-dark as the study type that proves the value of existing, live campaigns. A small dark cell or a budget-neutral multi-cell design limits the revenue exposed.

Free tools for this topic

FREE TOOLAttribution DoctorA media-mix model that runs in your browser.FREE TOOLUTM Campaign BuilderClean tracking links your analytics will thank you for.PLAYBOOKThe First-Party Data PlaybookMeasurement that survives privacy — and gets sharper.

Keep reading

GlossaryWhat Is Incrementality? The Question Attribution Can't AnswerRead →GuidesHow to Scale Paid Ads Without Killing ROASRead →GlossaryWhat Is ROAS? Formula, Benchmarks & Break-Even MathRead →
CATALIST NEWSLETTER

Monthly dose of growth marketing.

Get marketing tips, narratives, guides, and playbooks delivered to your inbox.

Protected by reCAPTCHA — Google's Privacy Policy and Terms of Service apply.