Incrementality

How to Run a Quarterly Geo-Holdout Test Your CFO Will Actually Trust

A step-by-step operator guide to quarterly geo-holdout incrementality tests: market matching, cell sizing, spend floors, contamination traps, and how to read a null result.

On this page

Every platform dashboard will tell you your ads work. A geo-holdout test is how you find out whether that's true, and 2026 is the year the question stopped being optional. Click-based attribution is losing its raw material: Pew Research Center found users clicked a traditional result on just 8% of searches that showed an AI summary, versus 15% without one, and Ahrefs measured a 58% drop in clicks to top-ranking pages when an AI Overview is present. Meanwhile agentic traffic doesn't carry a human's cookie, and auto-upgraded campaign types blur what a "click" even attributes to. The industry has noticed. Per eMarketer, 60% of US senior decision-makers now trust independent incrementality testing most among measurement approaches, versus 37% for in-platform reporting.

So here is the full procedure I run with clients each quarter: one channel, one question, one clean geographic experiment. It takes about ten weeks end to end. Follow the steps in order.

Step 1: Write the decision before the design

A geo test answers exactly one question, so pick the question a real dollar decision hangs on. Bad framing: "is Meta incremental?" Good framing: "if Advantage+ prospecting returns under 1.0 iROAS, we cut its budget 30% next quarter; above 2.0, we scale it 25%."

Write that sentence down, get the person who owns the budget to sign it, and define the metric now: total revenue (or new-customer revenue) in the test regions, from your order system, keyed on shipping address. Platform-reported conversions are the thing on trial; they cannot also be the judge. If you need a refresher on why causal lift and attributed credit are different animals, start with what incrementality actually measures.

Step 2: Match your markets

You need two groups of geographies that behave like twins before the test starts. In the US, work at DMA level (210 markets); elsewhere, use regions, provinces, or postcode clusters large enough to have stable weekly revenue.

The practical method in 2026 is synthetic control: instead of pairing Denver with Kansas City by eyeball, an algorithm builds a weighted blend of control markets that tracks each test market's history. Meta's GeoLift is open source, well documented, and does the matching, power analysis, and readout in one workflow. Feed it at least 8 weeks (ideally 12) of daily or weekly revenue by geo.

Two quality gates before you proceed:

  1. The synthetic control should track holdout markets with a mean error under roughly 5% per week in the pre-period. If your fit is worse, your markets are too small or too volatile; merge them into bigger clusters.
  2. Exclude weird geos entirely: markets with a retail door opening, a regional promo calendar, one dominant B2B account, or a history of stockouts. One anomalous market can swamp the readout.

Step 3: Size the cells and set the spend floor

This is where most DIY tests quietly fail. The math is unforgiving: you can only detect an effect that is larger than your noise.

  • Cell size. Holdout cells covering 20–30% of national revenue are the workable range for most brands. Below 15%, noise usually wins; above 35%, the revenue you're risking starts to hurt.
  • Minimum spend. The channel's spend inside the holdout geos, as a share of those geos' revenue, sets your ceiling. If you spend $10k/month in markets doing $1M/month, even a perfect 2.0 iROAS moves revenue 2%, which is invisible against 5% weekly noise. Rule of thumb: expected incremental revenue (claimed ROAS × withheld spend) should be at least twice your pre-period matching error, ideally 3x.
  • Duration. Six weeks of treatment is the floor for considered purchases; four can work for impulse categories with high order volume. Shorter tests confound lift with lag.

If the power analysis says you can't detect the effect you care about, do not run the test. Raise spend concentration in test geos first, or test a bigger decision.

Map with pins marking regional test and control markets
Photo by Z on Unsplash

Step 4: Lock the calendar and seal the borders

Contamination kills more geo tests than bad math does. Before launch, walk this checklist:

  1. Geo settings that actually hold. Set location targeting to physical presence, never "presence or interest." Verify that Performance Max and Advantage+ campaigns respect your exclusion lists; auto-upgraded campaign types have a habit of redrawing their own maps.
  2. Retargeting pools. Audiences built before the blackout contain holdout users. Exclude holdout geos from retargeting too, or your "dark" markets keep seeing ads.
  3. IP versus shipping. People browse from an office in one DMA and ship home to another. Measuring revenue by shipping address (Step 1) absorbs most of this, but expect a few percent of unavoidable bleed and note it in the writeup.
  4. Freeze everything else. No national promos, no PR pushes, no email calendar changes, no other channel "helpfully" backfilling soft markets. Smart bidding on your remaining channels will try to chase the demand gap; cap those budgets by geo if you can.
  5. Log confounders you can't control. Weather events, a competitor's regional push, retail distribution changes. You'll want the list when you read results.

Then go dark in the holdout cells and touch nothing for the full window.

A worked example

A DTC home goods brand, roughly $1M/month in revenue, spends $150k/month on Meta. The platform reports a 3.0 ROAS, which implies Meta drives $450k, or 45% of all revenue. The CFO is skeptical. Fair.

Quarterly geo-holdout design: Meta prospecting
ParameterValue
QuestionIs Meta prospecting iROAS above 1.0?
Decision ruleUnder 1.0: cut budget 30%. Over 2.0: scale 25%.
Pre-period8 weeks, daily revenue by DMA
Holdout cell17 of 68 DMAs, 25% of revenue (~$250k/mo baseline)
Withheld spend$37.5k/month (25% of Meta budget)
Matching error (MAPE)4.1% weekly, synthetic control via GeoLift
Treatment window6 weeks dark, 2-week cooldown
Minimum detectable effect~4.5% revenue gap at 90% confidence
Illustrative design based on a composite client engagement.

If the platform's 3.0 ROAS were causal, going dark should depress holdout revenue by about 45%. The minimum detectable effect is 4.5%, so the test is heavily overpowered for the claim; even an iROAS of 0.5 (a 7.5% gap) is detectable.

The readout: holdout revenue came in 9% below its synthetic control, with a 90% interval of 5–13%. That's roughly $22.5k/month of genuinely incremental revenue against $37.5k of withheld spend. Measured iROAS: 0.6, versus the reported 3.0. The decision rule fires: cut 30%, reallocate, re-test the smaller budget next quarter. Nobody argues, because everyone signed the rule in Step 1.

Step 5: Reading a null result honestly

Sometimes the gap is statistically zero. Read it in this order:

  1. Check power first. Was your minimum detectable effect smaller than the effect the platform claims? If the platform says 3.0 ROAS and your MDE only allowed you to see effects above 4.0, the test was uninformative. That's a design failure, and you say so plainly.
  2. Audit contamination. Pull delivery reports by geo. Any impressions in holdout markets above ~2–3% of normal levels weakens the read.
  3. Then believe it. If the test was powered and clean, a null is a finding: this channel, at this budget, in this period, did not produce measurable incremental revenue. Averaged over the range of true effects consistent with your interval, that spend is doing very little.

What a null does not tell you: that the channel can never work, that creative was good, or that a different funnel stage would fail too. It prices one configuration. That's still more than click attribution can honestly claim right now.

When the test says the channel isn't working

This is the part most guides skip. You will, at some point, get a result like the worked example, or worse, and it will implicate a channel someone's career is attached to.

Cut in stages rather than to zero. Take the budget down 25–40%, hold four weeks, and watch blended revenue and CAC. Big platforms have diminishing returns curves; a channel with 0.6 iROAS at $150k/month might clear 1.5 at $80k. Your test measured the average effect of the whole budget, and the marginal dollar is usually worse than the average one, which is exactly why scaling decisions need their own discipline.

Reallocate to the next test, and resist the urge to re-run the same experiment hoping for a friendlier answer. Per eMarketer, 52% of US brand and agency marketers already run incrementality tests and experiments; the ones getting value from them are the ones who let uncomfortable results change budgets.

Make it quarterly, and let it calibrate everything else

One test prices one channel in one quarter. The compounding value shows up when you rotate: Meta prospecting in Q1, branded search in Q2, TikTok in Q3, retargeting in Q4. After a year you have four causal anchor points, and those anchors are the best calibration data your media mix model will ever get; an MMM constrained by real experiments is a different instrument from one fit on observational data alone. If you're still deciding how experiments, MMM, and attribution fit together, the MTA versus MMM comparison covers the division of labor, and our media mix calculator helps you sketch the reallocation math before the next budget lock.

The quarterly rhythm also solves the political problem. When incrementality testing is a standing process rather than an ambush, results stop being accusations and start being weather reports.

Sources

Frequently asked questions

How long should a geo-holdout test run?
Plan roughly ten weeks end to end: at least eight weeks of pre-period data for market matching, six weeks of treatment, and a two-week cooldown to catch lagged conversions. Considered purchases with long research cycles need the full six weeks or more; high-volume impulse categories can sometimes read cleanly in four. Shorter windows confound the lift you want to measure with ordinary purchase lag, which biases results toward zero.
What if I only sell in one small country?
Work with whatever geographic units your ad platforms can target cleanly: regions, provinces, or postcode clusters. The constraint is statistical rather than political; each cell needs stable weekly revenue, so merge small areas until the synthetic control fits with under roughly 5% weekly error. Brands in small markets often need holdout cells covering 30% or more of revenue to reach adequate power, which raises the cost of testing but rarely makes it impossible.
Is a go-dark holdout better than a scale-up test?
They answer different questions. A go-dark holdout measures the average effect of your existing budget, which is what a CFO usually wants priced. A scale-up test measures the marginal return of the next dollar, which matters when you are deciding whether to grow a channel. If the channel is politically sensitive or the budget is large, run the holdout first; it is the harder, more decision-relevant number.
How is a geo test different from media mix modeling?
MMM is an observational model fit on historical spend and revenue; it estimates everything at once but inherits the correlations in your past behavior. A geo test is a true experiment on one channel, so it produces a causal number with fewer assumptions. They work best together: quarterly geo tests provide calibration anchors that constrain the MMM, and the MMM extends those anchors across channels and time periods you have not tested.

Free tools for this topic

FREE TOOLAI Brand Visibility MonitorDoes ChatGPT recommend you — or your competitor?CALCULATORROAS & Break-Even CalculatorKnow the ROAS you actually need before you scale.PLAYBOOKThe AI Search PlaybookGet cited by ChatGPT, Perplexity and Google AI Overviews.

Keep reading

GlossaryWhat Is Incrementality? The Question Attribution Can't AnswerRead →GlossaryWhat Is Media Mix Modeling? MMM, ExplainedRead →ComparisonsMulti-Touch Attribution vs Media Mix ModelingRead →
CATALIST NEWSLETTER

Monthly dose of growth marketing.

Get marketing tips, narratives, guides, and playbooks delivered to your inbox.