Docs

Geo-Incrementality Testing: Data to Track

Google's Meridian GeoX brings open-source geo-incrementality testing to marketing mix models — the in-app data that makes lift readable by region.

Acquisition

Google previewed Meridian GeoX this year, an open-source geo-incrementality testing tool that plugs directly into Meridian, its marketing mix modelling framework. The mechanic is simple to describe and hard to do well without the right data: hold out advertising in some regions, run it normally in others, and measure the difference in outcomes to isolate the actual incremental lift a channel caused rather than the lift it was merely credited with by last-touch attribution. GeoX adds native multi-cell execution, so several treatments can be compared in the same study, and feeds the resulting signals back into the broader MMM to sharpen budget allocation.

For app teams, geo-incrementality testing only works if in-app outcome data can be sliced cleanly by region and time window in the first place. A geo experiment is only as trustworthy as the app-side metric it’s measuring against — if install volume, activation, or revenue can’t be reliably attributed to a user’s region at the time of the experiment, or if regional data is too sparse to reach significance within a reasonable test window, the “incremental lift” a geo test reports is noise wearing the shape of a causal finding. Teams used to channel-level attribution dashboards often don’t have region-level outcome data broken out cleanly enough to support this kind of test, because nothing before now demanded it.

Data Points to Track

  • Region assigned at install (or session) time, captured as a stable property rather than inferred later from IP or billing address, so experiment and control regions can be compared on consistent footing
  • Core outcome metrics segmented by region and time window, matching whatever windows a geo test defines, for installs, activation, and any revenue event the test is meant to influence
  • Baseline regional variance before any test starts, since regions differ in size and behaviour even without any experiment running, and that baseline noise has to be understood before a lift number can be trusted
  • Cross-region contamination signals, such as users travelling between test and control regions, or campaigns that weren’t actually geo-fenced as intended, which quietly undermine a geo test’s validity
  • Time-to-significance per metric, tracked per test, so low-volume regions or metrics that would need months to reach a reliable result aren’t reported with false confidence at week two

Setup Steps

  1. Confirm region is captured reliably at the point of first touch, not backfilled from a billing address weeks later, since geo tests depend on knowing where a user was when the test was live.
  2. Establish a regional baseline for each core outcome metric over a normal period before running any test, so the “no ads” control regions have a known reference point rather than an assumed flat line.
  3. Build a dashboard that can slice install, activation, and revenue by region and date range on demand, rather than requiring a one-off data pull every time marketing wants to check a geo test.
  4. Flag likely contamination sources — users who travel, shared devices, or campaigns that leak outside their intended geo-fence — before trusting a test’s headline lift number.
  5. Set a minimum sample and duration threshold per metric and region size before a geo test’s result is treated as final, rather than acting on early, statistically thin reads.

Actionable Insights

A geo test result should always be read against the pre-test regional baseline, not treated as a standalone number — a “20% lift” in a region with historically volatile week-to-week install numbers means something very different from the same lift in a stable region. Where in-app revenue or activation data can’t yet be reliably sliced by region, that’s the gap to close before running a geo test at all, since the marketing team’s confidence in the result will only ever be as strong as the weakest link in the outcome data feeding it. Once regional data quality is solid, geo-incrementality results are worth weighting more heavily than last-touch attribution in budget conversations — they’re measuring causation, which last-touch models were never built to do.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert