How to Calibrate a Marketing Mix Model With Geo Lift Test Results
One geo lift readout worked through Google Meridian and Meta Robyn: the ROI prior, the calibration row, and how to match KPI, channel scope and time window.
On this page
To calibrate a marketing mix model using geo-based incrementality test results, convert the readout into the one input your modeling tool accepts, then confirm that input describes the same revenue, the same channel and the same weeks the model sees. Google's Meridian takes the test's ROI point estimate and standard error, with spend and dates, and builds a prior on the channel's ROI. Meta's Robyn takes an absolute lift for a channel and date range and scores every candidate model on its distance from that number. We worked one hypothetical test through both tools, using each one's documentation as it read on October 7, 2026.
The geo lift readout we start from
Every figure in this example is hypothetical. A retailer models weekly revenue over 104 weeks ending September 13, 2026, and paid social averages $5,000 a day across that window. From May 4 to June 14, 2026, the team runs a geo holdout test in which paid social goes dark in markets that normally take 30% of the channel's spend.
| Readout item | Hypothetical value |
|---|---|
| Spend withheld in holdout markets over six weeks | $63,000 |
| Paid social spend that still ran in other markets | $147,000 |
| Revenue lost against the synthetic control | $157,500 |
| ROI (lost revenue ÷ withheld spend) | 2.5 |
| 95% interval on ROI | 1.5 to 3.5 |
| P-value | 0.01 |
Our guide to analyzing an incrementality test result covers how to read those rows. What matters for calibration is that the two tools take different ones. Meridian needs a standard error, and a geo readout may not print it: the summary in the GeoLift walkthrough from Meta shows Average ATT, Percent Lift, Incremental Y and a P-value, and we found no standard error in that output. Our workaround assumes a symmetric, roughly normal interval and divides half its width by 1.96, which gives 1.0 ÷ 1.96, or 0.51.
How does the test become a Meridian ROI prior?
MMM calibration with a geo experiment or lift test starts, in Meridian, with a prior. The calibration guide from Google defines the term in one line: "Calibration is the process of using experiment results and other domain knowledge to set channel-specific priors." The same page cautions that "There is no single formula to translate an experiment result into a prior."
Meridian's docs describe a CalibrationBuilder that automates the translation. Results from Meridian GeoX pass through a dedicated method, and any other test, ours included, is registered with with_incrementality_experiment_result. The API reference from Google lists the arguments, and our readout fills them as follows.
builder = prior_builder.CalibrationBuilder(data)
builder.with_incrementality_experiment_result(
'paid_social',
point_estimates=2.5,
standard_errors=0.51,
experiment_kpi_types='revenue',
experiment_total_spends=63000,
experiment_start_dates='2026-05-04',
experiment_end_dates='2026-06-14',
)
calibrated = builder.build()
build() returns calibrated roi_m and roi_rf distributions, which go into PriorDistribution and then ModelSpec. The reference describes the spend argument as "the typical total spend for the experiment duration for the channel", and we read that as the $63,000 the holdout markets would normally have received. It also lists the conditions that raise an error, among them a channel name missing from the model's paid channels, experiment spend above the channel's total spend, and a statistically significantly negative lift.
Why does the prior come out wider than the test's interval?
Geo experiments calibrate marketing mix models most directly through Bayesian calibration, where the geo lift's incrementality estimate enters the MMM as a prior distribution, and Meridian widens that distribution on purpose. The builder's guide from Google multiplies the experiment's standard error by the square root of one plus adjustments for spend, recency and duration, and it scales the point estimate for duration. Using the documented defaults (geometric decay, alpha of 0.5, max lag of 8, a 52-week recency half-life), our arithmetic for this test runs as follows.
| Adjustment | What the test supplies | Result |
|---|---|---|
| Spend | $1,500 a day in the test against $5,000 a day for the channel, a ratio of 0.3 | Adds 2.33 under the square root |
| Recency | The test ended 13 weeks before the last modeled date | Adds 0.19 |
| Duration | Six weeks capture 98.6% of the adstock weights | Adds 0.014 and scales ROI by 1.014 |
| Adjusted point estimate | 2.5 × 1.014 | 2.53 |
| Adjusted standard error | 0.51 × √3.54 | 0.96 |
By this hand arithmetic, which counts the six-week test as 42 days, a readout of 2.5 with a standard error of 0.51 turns into a prior centered near 2.53 with a standard error of 0.96, nearly twice as wide. The package can land slightly lower: the calibration source from Google divides test spend by the end date minus the start date, 41 days for these dates, which by our arithmetic gives an adjusted standard error near 0.95. Almost all of that widening comes from the spend term, because the holdout covered markets carrying 30% of the channel's spend. With the default baseline prior, the builder then fits a LogNormal or Gamma distribution to the result. Teams building a prior by hand can use lognormal_dist_from_mean_std, a helper the calibration guide documents.
A note in the builder's guide limits which tests belong here. The builder is designed for experiments measured against "a counterfactual baseline of zero spend", which a go-dark holdout satisfies. Meridian does not block results from experiments with non-zero counterfactuals, and its example is marginal lift measured against reduced spend, but the note says doing so "is not recommended" because of the estimand mismatch with the MMM's zero-spend counterfactual. By the same definition, a heavy-up test that compares extra spend with normal spend falls on the not-recommended side.
What does Robyn's calibration input expect?
Meta's Robyn marketing mix modeling documentation treats calibration with experiments as a third objective for its optimizer. The features page from Meta puts it this way: "Robyn implements the MMM calibration as an objective function in the multi-objective optimization by parameterizing the difference between causal results and predicted media contribution." That calibration error is named MAPE.LIFT, and it sits beside the prediction error (NRMSE) and the business error (DECOMP.RSSD).
The input is a data frame handed to robyn_inputs(). The package's demo script, also from Meta, shows the eight columns, and our test fills them as follows.
calibration_input <- data.frame(
channel = "paid_social_S",
liftStartDate = as.Date("2026-05-04"),
liftEndDate = as.Date("2026-06-14"),
liftAbs = 367500,
spend = 147000,
confidence = 0.99,
metric = "revenue",
calibration_scope = "immediate"
)
liftAbs and spend are the fields that take thought. The demo's guide says Robyn currently accepts only a point estimate, and that the estimate has to match the spend in the model's variable. Its example is a channel that usually spends $100K a week under a 70% holdout, which should be entered with the estimate for the $30K. Our national paid social variable holds the $147,000 that still ran during the test, so that figure goes in spend, and we apply the measured ROI of 2.5 to it for a liftAbs of $367,500. That step rests on our own assumption that the markets left on return what the holdout markets did, which the test did not measure.
For confidence, the demo suggests 1 minus the p-value for a frequentist experiment, and it tells users to set calibration_scope to immediate when calibrating with causal experiments. The interval never reaches the model. In the lift calibration code that Meta keeps on the repository's main branch, the error is the absolute gap between liftAbs and the model's predicted effect for the channel over the lift dates (less carryover from earlier spend under the immediate scope), divided by liftAbs. We found no line in that function that weights a study by its confidence value.
Does the input describe what the test measured?
In marketing mix modeling, geo experiments only improve calibration when the causal lift and the model's incrementality estimate describe the same outcome, the same media and the same weeks. Each tool enforces part of that match and leaves the remainder to the analyst.
| What must match | Meridian | Robyn |
|---|---|---|
| KPI | experiment_kpi_types is 'revenue' or 'non-revenue'. Non-revenue results are converted with the model's revenue_per_kpi, and the call errors without it | metric must equal dep_var, or the run stops |
| Channel scope | The channel must be one of the model's paid channels. A campaign-level test against a channel-level model is handled through the spend ratio, and a small campaign widens the prior more | channel must be a paid media or organic variable. A test across several is entered as channel_A+channel_B |
| Spend | Experiment spend cannot exceed the channel's total spend | A warning fires when spend is more than 10% away from the model's spend for those dates |
| Time window | The dates drive the recency and duration adjustments | Lift dates must fall inside the modeling window, or the run stops |
| Uncertainty | The standard error sets the prior's width | Point estimate only. Confidence under 0.8 draws a warning |
The Robyn column reflects the input checks that Meta ships in the package source. Its spend warning is the one most likely to fire on a geo test, because a national model's spend for the test dates rarely equals the spend in the tested markets. The warning text suggests rechecking the dates or splitting the media input into separate channels.
On the Meridian side, the roi_calibration argument can tie a prior to a specific time period, yet the calibration guide recommends setting the prior on the full modeling time window whenever possible. The builder's adjustments point the same way: an older or shorter test gets a wider prior, and the window stays whole.
What to check once the calibrated model has run
For Meridian, the calibration guide recommends plotting custom priors, and it notes that comparing prior with posterior can show whether the model is learning from the data or being strongly influenced by the prior. For Robyn, each candidate model's calibration error reports how close its paid social effect over those six weeks landed to $367,500.
The comparison also bears on test design. By our arithmetic from Meridian's formula, the same readout from markets carrying 60% of spend would have produced an adjusted standard error of 0.70 where ours produced 0.96. Robyn would treat both tests alike once the point estimate and the spend lined up. A team planning its next holdout around a Bayesian model has a reason to weigh market coverage alongside test length.
The same readout has a second use in correcting platform-reported conversions through the multiplier method, and our primer on what incrementality means explains why a lift test carries this much weight. Attribution and MMM work sits with our data and analytics team, and the Attribution Doctor is a further resource on the attribution side.
Sources
- Google, Meridian guide "Calibrate treatment priors", last updated September 24, 2026
- Google, Meridian guide "Set custom ROI priors using past experiments", last updated September 23, 2026
- Google, Meridian API reference for CalibrationBuilder, last updated September 24, 2026
- Google, Meridian calibration source, main branch
- Meta, Robyn documentation, "Key Features"
- Meta, Robyn demo script, main branch
- Meta, Robyn calibration source, main branch
- Meta, Robyn input checks, main branch
- Meta, GeoLift walkthrough
