Most teams measure experiments one at a time: ship a change, compare the test group to a control group, declare a winner if the lift clears significance. What that approach can’t answer is a bigger question — is the experimentation program as a whole actually growing the business, or is it a string of small individual “wins” that cancel each other out, regress after launch, or double-count the same underlying users across overlapping tests? A global holdout group — a small, stable slice of users permanently excluded from every test and kept on the baseline experience — answers that question directly, by giving you a clean comparison point for the aggregate effect of everything you’ve shipped.
The catch is that a holdout group is only as good as the tracking discipline around it. If holdout users leak into individual experiments, if the holdout isn’t re-randomised when its composition drifts, or if the metrics compared between holdout and non-holdout aren’t the same ones used in individual experiment reports, the number you get back looks authoritative but means nothing. A holdout is a long-running measurement instrument, not a one-off test, and it needs the same rigor as any other piece of production infrastructure — assignment integrity, contamination checks, and a stable metric definition that doesn’t shift every quarter.
Data Points to Track
- Holdout assignment ID per user, persisted independently of any individual experiment’s variant assignment, so holdout membership can be audited separately from test participation
- Contamination events — cases where a holdout user was accidentally exposed to a shipped experiment variant — logged with enough detail to exclude or correct for the leak
- Metric parity between holdout reporting and individual experiment reporting, so the same conversion or revenue metric isn’t defined one way in the holdout dashboard and another way in per-test results
- Holdout cohort size and churn over time, since a holdout that shrinks or refreshes unpredictably undermines the statistical power of the comparison
- Cumulative feature-ship timeline overlaid on the holdout comparison window, making it possible to attribute a change in the holdout gap to a specific batch of releases
Setup Steps
- Assign holdout membership at the same layer as experiment randomisation, using a dedicated flag that experiment targeting logic explicitly excludes, rather than relying on a naming convention or manual list.
- Add an automated contamination check that flags any holdout user who received a non-default variant in any active experiment, and alert on it rather than discovering it at reporting time.
- Standardise the core business metrics (activation, retention, revenue per user) used in holdout comparisons so they match the definitions used in individual experiment scorecards.
- Set a fixed re-evaluation cadence for holdout size and composition — quarterly is common — rather than letting the cohort drift indefinitely.
- Log every feature or experiment that ships to the non-holdout population, timestamped, so a change in the holdout gap can be traced back to a specific release window.
Actionable Insights
The holdout gap — the difference in your core metric between the holdout group and everyone else — is the closest thing you’ll get to a true read on whether your experimentation program is net-positive. A shrinking or flat gap despite a steady stream of “winning” experiments is a signal that individual test results aren’t translating into real aggregate impact, often because of novelty effects, metric gaming, or interaction effects between overlapping tests. Treat holdout tracking as a program-level audit function sitting alongside individual experiment analysis, not a replacement for it — the two answer different questions, and you need both to trust either one.
Related Resources
Need help tracking this in your app?
Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.
Talk to an expert