How to Run a Quarterly Geo-Holdout Test Your CFO Will Actually Trust
A step-by-step operator guide to quarterly geo-holdout incrementality tests: market matching, cell sizing, spend floors, contamination traps, and how to read a null result.
On this page
- Step 1: Write the decision before the design
- Step 2: Match your markets
- Step 3: Size the cells and set the spend floor
- Step 4: Lock the calendar and seal the borders
- A worked example
- Step 5: Reading a null result honestly
- When the test says the channel isn't working
- Make it quarterly, and let it calibrate everything else
- Sources
Every platform dashboard will tell you your ads work. A geo-holdout test is how you find out whether that's true, and 2026 is the year the question stopped being optional. Click-based attribution is losing its raw material: Pew Research Center found users clicked a traditional result on just 8% of searches that showed an AI summary, versus 15% without one, and Ahrefs measured a 58% drop in clicks to top-ranking pages when an AI Overview is present. Meanwhile agentic traffic doesn't carry a human's cookie, and auto-upgraded campaign types blur what a "click" even attributes to. The industry has noticed. Per eMarketer, 60% of US senior decision-makers now trust independent incrementality testing most among measurement approaches, versus 37% for in-platform reporting.
So here is the full procedure I run with clients each quarter: one channel, one question, one clean geographic experiment. It takes about ten weeks end to end. Follow the steps in order.
Step 1: Write the decision before the design
A geo test answers exactly one question, so pick the question a real dollar decision hangs on. Bad framing: "is Meta incremental?" Good framing: "if Advantage+ prospecting returns under 1.0 iROAS, we cut its budget 30% next quarter; above 2.0, we scale it 25%."
Write that sentence down, get the person who owns the budget to sign it, and define the metric now: total revenue (or new-customer revenue) in the test regions, from your order system, keyed on shipping address. Platform-reported conversions are the thing on trial; they cannot also be the judge. If you need a refresher on why causal lift and attributed credit are different animals, start with what incrementality actually measures.
Step 2: Match your markets
You need two groups of geographies that behave like twins before the test starts. In the US, work at DMA level (210 markets); elsewhere, use regions, provinces, or postcode clusters large enough to have stable weekly revenue.
The practical method in 2026 is synthetic control: instead of pairing Denver with Kansas City by eyeball, an algorithm builds a weighted blend of control markets that tracks each test market's history. Meta's GeoLift is open source, well documented, and does the matching, power analysis, and readout in one workflow. Feed it at least 8 weeks (ideally 12) of daily or weekly revenue by geo.
Two quality gates before you proceed:
- The synthetic control should track holdout markets with a mean error under roughly 5% per week in the pre-period. If your fit is worse, your markets are too small or too volatile; merge them into bigger clusters.
- Exclude weird geos entirely: markets with a retail door opening, a regional promo calendar, one dominant B2B account, or a history of stockouts. One anomalous market can swamp the readout.
Step 3: Size the cells and set the spend floor
This is where most DIY tests quietly fail. The math is unforgiving: you can only detect an effect that is larger than your noise.
- Cell size. Holdout cells covering 20–30% of national revenue are the workable range for most brands. Below 15%, noise usually wins; above 35%, the revenue you're risking starts to hurt.
- Minimum spend. The channel's spend inside the holdout geos, as a share of those geos' revenue, sets your ceiling. If you spend $10k/month in markets doing $1M/month, even a perfect 2.0 iROAS moves revenue 2%, which is invisible against 5% weekly noise. Rule of thumb: expected incremental revenue (claimed ROAS × withheld spend) should be at least twice your pre-period matching error, ideally 3x.
- Duration. Six weeks of treatment is the floor for considered purchases; four can work for impulse categories with high order volume. Shorter tests confound lift with lag.
If the power analysis says you can't detect the effect you care about, do not run the test. Raise spend concentration in test geos first, or test a bigger decision.
Step 4: Lock the calendar and seal the borders
Contamination kills more geo tests than bad math does. Before launch, walk this checklist:
- Geo settings that actually hold. Set location targeting to physical presence, never "presence or interest." Verify that Performance Max and Advantage+ campaigns respect your exclusion lists; auto-upgraded campaign types have a habit of redrawing their own maps.
- Retargeting pools. Audiences built before the blackout contain holdout users. Exclude holdout geos from retargeting too, or your "dark" markets keep seeing ads.
- IP versus shipping. People browse from an office in one DMA and ship home to another. Measuring revenue by shipping address (Step 1) absorbs most of this, but expect a few percent of unavoidable bleed and note it in the writeup.
- Freeze everything else. No national promos, no PR pushes, no email calendar changes, no other channel "helpfully" backfilling soft markets. Smart bidding on your remaining channels will try to chase the demand gap; cap those budgets by geo if you can.
- Log confounders you can't control. Weather events, a competitor's regional push, retail distribution changes. You'll want the list when you read results.
Then go dark in the holdout cells and touch nothing for the full window.
A worked example
A DTC home goods brand, roughly $1M/month in revenue, spends $150k/month on Meta. The platform reports a 3.0 ROAS, which implies Meta drives $450k, or 45% of all revenue. The CFO is skeptical. Fair.
| Parameter | Value |
|---|---|
| Question | Is Meta prospecting iROAS above 1.0? |
| Decision rule | Under 1.0: cut budget 30%. Over 2.0: scale 25%. |
| Pre-period | 8 weeks, daily revenue by DMA |
| Holdout cell | 17 of 68 DMAs, 25% of revenue (~$250k/mo baseline) |
| Withheld spend | $37.5k/month (25% of Meta budget) |
| Matching error (MAPE) | 4.1% weekly, synthetic control via GeoLift |
| Treatment window | 6 weeks dark, 2-week cooldown |
| Minimum detectable effect | ~4.5% revenue gap at 90% confidence |
If the platform's 3.0 ROAS were causal, going dark should depress holdout revenue by about 45%. The minimum detectable effect is 4.5%, so the test is heavily overpowered for the claim; even an iROAS of 0.5 (a 7.5% gap) is detectable.
The readout: holdout revenue came in 9% below its synthetic control, with a 90% interval of 5–13%. That's roughly $22.5k/month of genuinely incremental revenue against $37.5k of withheld spend. Measured iROAS: 0.6, versus the reported 3.0. The decision rule fires: cut 30%, reallocate, re-test the smaller budget next quarter. Nobody argues, because everyone signed the rule in Step 1.
Step 5: Reading a null result honestly
Sometimes the gap is statistically zero. Read it in this order:
- Check power first. Was your minimum detectable effect smaller than the effect the platform claims? If the platform says 3.0 ROAS and your MDE only allowed you to see effects above 4.0, the test was uninformative. That's a design failure, and you say so plainly.
- Audit contamination. Pull delivery reports by geo. Any impressions in holdout markets above ~2–3% of normal levels weakens the read.
- Then believe it. If the test was powered and clean, a null is a finding: this channel, at this budget, in this period, did not produce measurable incremental revenue. Averaged over the range of true effects consistent with your interval, that spend is doing very little.
What a null does not tell you: that the channel can never work, that creative was good, or that a different funnel stage would fail too. It prices one configuration. That's still more than click attribution can honestly claim right now.
When the test says the channel isn't working
This is the part most guides skip. You will, at some point, get a result like the worked example, or worse, and it will implicate a channel someone's career is attached to.
Cut in stages rather than to zero. Take the budget down 25–40%, hold four weeks, and watch blended revenue and CAC. Big platforms have diminishing returns curves; a channel with 0.6 iROAS at $150k/month might clear 1.5 at $80k. Your test measured the average effect of the whole budget, and the marginal dollar is usually worse than the average one, which is exactly why scaling decisions need their own discipline.
Reallocate to the next test, and resist the urge to re-run the same experiment hoping for a friendlier answer. Per eMarketer, 52% of US brand and agency marketers already run incrementality tests and experiments; the ones getting value from them are the ones who let uncomfortable results change budgets.
Make it quarterly, and let it calibrate everything else
One test prices one channel in one quarter. The compounding value shows up when you rotate: Meta prospecting in Q1, branded search in Q2, TikTok in Q3, retargeting in Q4. After a year you have four causal anchor points, and those anchors are the best calibration data your media mix model will ever get; an MMM constrained by real experiments is a different instrument from one fit on observational data alone. If you're still deciding how experiments, MMM, and attribution fit together, the MTA versus MMM comparison covers the division of labor, and our media mix calculator helps you sketch the reallocation math before the next budget lock.
The quarterly rhythm also solves the political problem. When incrementality testing is a standing process rather than an ambush, results stop being accusations and start being weather reports.
Sources
- Pew Research Center: Google users are less likely to click on links when an AI summary appears
- Ahrefs: AI Overviews reduce clicks by 58%
- eMarketer: Incrementality testing earns marketers' top trust
- eMarketer: MMM, incrementality, and other measurement trends that will define 2026
- Meta GeoLift (open source)
