MMM Calibration

How to Calibrate a Marketing Mix Model With Geo Lift Test Results

One geo lift readout worked through Google Meridian and Meta Robyn: the ROI prior, the calibration row, and how to match KPI, channel scope and time window.

On this page

To calibrate a marketing mix model using geo-based incrementality test results, convert the readout into the one input your modeling tool accepts, then confirm that input describes the same revenue, the same channel and the same weeks the model sees. Google's Meridian takes the test's ROI point estimate and standard error, with spend and dates, and builds a prior on the channel's ROI. Meta's Robyn takes an absolute lift for a channel and date range and scores every candidate model on its distance from that number. We worked one hypothetical test through both tools, using each one's documentation as it read on October 7, 2026.

The geo lift readout we start from

Every figure in this example is hypothetical. A retailer models weekly revenue over 104 weeks ending September 13, 2026, and paid social averages $5,000 a day across that window. From May 4 to June 14, 2026, the team runs a geo holdout test in which paid social goes dark in markets that normally take 30% of the channel's spend.

Readout itemHypothetical value
Spend withheld in holdout markets over six weeks$63,000
Paid social spend that still ran in other markets$147,000
Revenue lost against the synthetic control$157,500
ROI (lost revenue ÷ withheld spend)2.5
95% interval on ROI1.5 to 3.5
P-value0.01

Our guide to analyzing an incrementality test result covers how to read those rows. What matters for calibration is that the two tools take different ones. Meridian needs a standard error, and a geo readout may not print it: the summary in the GeoLift walkthrough from Meta shows Average ATT, Percent Lift, Incremental Y and a P-value, and we found no standard error in that output. Our workaround assumes a symmetric, roughly normal interval and divides half its width by 1.96, which gives 1.0 ÷ 1.96, or 0.51.

How does the test become a Meridian ROI prior?

MMM calibration with a geo experiment or lift test starts, in Meridian, with a prior. The calibration guide from Google defines the term in one line: "Calibration is the process of using experiment results and other domain knowledge to set channel-specific priors." The same page cautions that "There is no single formula to translate an experiment result into a prior."

Meridian's docs describe a CalibrationBuilder that automates the translation. Results from Meridian GeoX pass through a dedicated method, and any other test, ours included, is registered with with_incrementality_experiment_result. The API reference from Google lists the arguments, and our readout fills them as follows.

builder = prior_builder.CalibrationBuilder(data)
builder.with_incrementality_experiment_result(
    'paid_social',
    point_estimates=2.5,
    standard_errors=0.51,
    experiment_kpi_types='revenue',
    experiment_total_spends=63000,
    experiment_start_dates='2026-05-04',
    experiment_end_dates='2026-06-14',
)
calibrated = builder.build()

build() returns calibrated roi_m and roi_rf distributions, which go into PriorDistribution and then ModelSpec. The reference describes the spend argument as "the typical total spend for the experiment duration for the channel", and we read that as the $63,000 the holdout markets would normally have received. It also lists the conditions that raise an error, among them a channel name missing from the model's paid channels, experiment spend above the channel's total spend, and a statistically significantly negative lift.

Why does the prior come out wider than the test's interval?

Geo experiments calibrate marketing mix models most directly through Bayesian calibration, where the geo lift's incrementality estimate enters the MMM as a prior distribution, and Meridian widens that distribution on purpose. The builder's guide from Google multiplies the experiment's standard error by the square root of one plus adjustments for spend, recency and duration, and it scales the point estimate for duration. Using the documented defaults (geometric decay, alpha of 0.5, max lag of 8, a 52-week recency half-life), our arithmetic for this test runs as follows.

AdjustmentWhat the test suppliesResult
Spend$1,500 a day in the test against $5,000 a day for the channel, a ratio of 0.3Adds 2.33 under the square root
RecencyThe test ended 13 weeks before the last modeled dateAdds 0.19
DurationSix weeks capture 98.6% of the adstock weightsAdds 0.014 and scales ROI by 1.014
Adjusted point estimate2.5 × 1.0142.53
Adjusted standard error0.51 × √3.540.96

By this hand arithmetic, which counts the six-week test as 42 days, a readout of 2.5 with a standard error of 0.51 turns into a prior centered near 2.53 with a standard error of 0.96, nearly twice as wide. The package can land slightly lower: the calibration source from Google divides test spend by the end date minus the start date, 41 days for these dates, which by our arithmetic gives an adjusted standard error near 0.95. Almost all of that widening comes from the spend term, because the holdout covered markets carrying 30% of the channel's spend. With the default baseline prior, the builder then fits a LogNormal or Gamma distribution to the result. Teams building a prior by hand can use lognormal_dist_from_mean_std, a helper the calibration guide documents.

A note in the builder's guide limits which tests belong here. The builder is designed for experiments measured against "a counterfactual baseline of zero spend", which a go-dark holdout satisfies. Meridian does not block results from experiments with non-zero counterfactuals, and its example is marginal lift measured against reduced spend, but the note says doing so "is not recommended" because of the estimand mismatch with the MMM's zero-spend counterfactual. By the same definition, a heavy-up test that compares extra spend with normal spend falls on the not-recommended side.

What does Robyn's calibration input expect?

Meta's Robyn marketing mix modeling documentation treats calibration with experiments as a third objective for its optimizer. The features page from Meta puts it this way: "Robyn implements the MMM calibration as an objective function in the multi-objective optimization by parameterizing the difference between causal results and predicted media contribution." That calibration error is named MAPE.LIFT, and it sits beside the prediction error (NRMSE) and the business error (DECOMP.RSSD).

The input is a data frame handed to robyn_inputs(). The package's demo script, also from Meta, shows the eight columns, and our test fills them as follows.

calibration_input <- data.frame(
  channel = "paid_social_S",
  liftStartDate = as.Date("2026-05-04"),
  liftEndDate = as.Date("2026-06-14"),
  liftAbs = 367500,
  spend = 147000,
  confidence = 0.99,
  metric = "revenue",
  calibration_scope = "immediate"
)

liftAbs and spend are the fields that take thought. The demo's guide says Robyn currently accepts only a point estimate, and that the estimate has to match the spend in the model's variable. Its example is a channel that usually spends $100K a week under a 70% holdout, which should be entered with the estimate for the $30K. Our national paid social variable holds the $147,000 that still ran during the test, so that figure goes in spend, and we apply the measured ROI of 2.5 to it for a liftAbs of $367,500. That step rests on our own assumption that the markets left on return what the holdout markets did, which the test did not measure.

For confidence, the demo suggests 1 minus the p-value for a frequentist experiment, and it tells users to set calibration_scope to immediate when calibrating with causal experiments. The interval never reaches the model. In the lift calibration code that Meta keeps on the repository's main branch, the error is the absolute gap between liftAbs and the model's predicted effect for the channel over the lift dates (less carryover from earlier spend under the immediate scope), divided by liftAbs. We found no line in that function that weights a study by its confidence value.

Does the input describe what the test measured?

In marketing mix modeling, geo experiments only improve calibration when the causal lift and the model's incrementality estimate describe the same outcome, the same media and the same weeks. Each tool enforces part of that match and leaves the remainder to the analyst.

What must matchMeridianRobyn
KPIexperiment_kpi_types is 'revenue' or 'non-revenue'. Non-revenue results are converted with the model's revenue_per_kpi, and the call errors without itmetric must equal dep_var, or the run stops
Channel scopeThe channel must be one of the model's paid channels. A campaign-level test against a channel-level model is handled through the spend ratio, and a small campaign widens the prior morechannel must be a paid media or organic variable. A test across several is entered as channel_A+channel_B
SpendExperiment spend cannot exceed the channel's total spendA warning fires when spend is more than 10% away from the model's spend for those dates
Time windowThe dates drive the recency and duration adjustmentsLift dates must fall inside the modeling window, or the run stops
UncertaintyThe standard error sets the prior's widthPoint estimate only. Confidence under 0.8 draws a warning

The Robyn column reflects the input checks that Meta ships in the package source. Its spend warning is the one most likely to fire on a geo test, because a national model's spend for the test dates rarely equals the spend in the tested markets. The warning text suggests rechecking the dates or splitting the media input into separate channels.

On the Meridian side, the roi_calibration argument can tie a prior to a specific time period, yet the calibration guide recommends setting the prior on the full modeling time window whenever possible. The builder's adjustments point the same way: an older or shorter test gets a wider prior, and the window stays whole.

What to check once the calibrated model has run

For Meridian, the calibration guide recommends plotting custom priors, and it notes that comparing prior with posterior can show whether the model is learning from the data or being strongly influenced by the prior. For Robyn, each candidate model's calibration error reports how close its paid social effect over those six weeks landed to $367,500.

The comparison also bears on test design. By our arithmetic from Meridian's formula, the same readout from markets carrying 60% of spend would have produced an adjusted standard error of 0.70 where ours produced 0.96. Robyn would treat both tests alike once the point estimate and the spend lined up. A team planning its next holdout around a Bayesian model has a reason to weigh market coverage alongside test length.

The same readout has a second use in correcting platform-reported conversions through the multiplier method, and our primer on what incrementality means explains why a lift test carries this much weight. Attribution and MMM work sits with our data and analytics team, and the Attribution Doctor is a further resource on the attribution side.

Sources

  • Google, Meridian guide "Calibrate treatment priors", last updated September 24, 2026
  • Google, Meridian guide "Set custom ROI priors using past experiments", last updated September 23, 2026
  • Google, Meridian API reference for CalibrationBuilder, last updated September 24, 2026
  • Google, Meridian calibration source, main branch
  • Meta, Robyn documentation, "Key Features"
  • Meta, Robyn demo script, main branch
  • Meta, Robyn calibration source, main branch
  • Meta, Robyn input checks, main branch
  • Meta, GeoLift walkthrough

Frequently asked questions

Do I need a standard error to calibrate Meridian with a geo lift test?
Yes, for a test run outside Meridian GeoX. The generic method on CalibrationBuilder takes a point estimate and a standard error for each experiment, and its reference lists a non-positive standard error among the conditions that raise an error. When a readout gives only an interval, the standard error has to be derived. We divided half the width of a 95% interval by 1.96, an approximation that assumes a symmetric, roughly normal interval.
Does Robyn use the confidence interval from a geo test?
Robyn's guide says it currently accepts only a point estimate as calibration input. The confidence column takes a single value, and the demo script suggests 1 minus the p-value for a frequentist experiment. The package's input checks warn when that value is under 0.8. In the lift calibration function we read on October 7, 2026, we found no line that weights a study by its confidence.
Can a heavy-up geo test calibrate a marketing mix model?
A heavy-up test compares extra spend with normal spend, so its counterfactual sits above zero. Meridian's documentation says CalibrationBuilder is designed for lift experiments measured against a counterfactual baseline of zero spend. It does not block results from tests with a non-zero counterfactual, giving marginal lift against reduced spend as its example, but it says doing so is not recommended because marginal lift introduces an estimand mismatch with the model's zero-spend counterfactual. A go-dark holdout matches the zero-spend definition.
What if we have several tests on the same channel?
Both tools accept more than one. Meridian's guide says that when multiple experiments are registered for a single channel, the builder combines them, along with an optional baseline prior, using Bayesian updating, after adjusting each one for spend, recency and duration. Robyn's calibration input is a data frame with one row per experiment, and its demo script shows four rows, including one that covers two channels at once.
Does the test have to fall inside the model's date range?
In Robyn it does: the input check stops the run when a lift's start or end date falls outside the modeling window. Meridian's builder uses the dates differently. It widens the prior for a test that ended before the last modeled date, using a 52-week half-life, and applies no recency adjustment when the last modeled date falls before the experiment. Google's guide still asks whether results from a different period apply.

Free tools for this topic

FREE TOOLAttribution DoctorA media-mix model that runs in your browser.FREE TOOLUTM Campaign BuilderClean tracking links your analytics will thank you for.PLAYBOOKThe First-Party Data PlaybookMeasurement that survives privacy — and gets sharper.

Keep reading

DataHow to Actually Analyze an Incrementality Test ResultRead →GuidesHow to Run a Quarterly Geo-Holdout Test Your CFO Will Actually TrustRead →DataThe Multiplier Method for Calibrating Platform Attribution With Incrementality TestsRead →
CATALIST NEWSLETTER

Monthly dose of growth marketing.

Get marketing tips, narratives, guides, and playbooks delivered to your inbox.

Protected by reCAPTCHA — Google's Privacy Policy and Terms of Service apply.