Docs

Firebase Remote Config A/B Testing: What to Track

Firebase has folded A/B Testing natively into Remote Config with usage-based pricing — the experiment data worth tracking before it skews.

Engagement

Firebase has merged A/B Testing directly into Remote Config, alongside Personalization and Rollouts, and moved Remote Config itself to usage-based pricing from September 2026 — a no-cost daily fetch tier, then pay-as-you-go on Blaze beyond it. For teams already using Remote Config to gate features and push copy changes, that’s convenient: experiments, personalised parameters, and staged rollouts now live in one surface instead of three. It also means three very different things — a plain config push, a personalisation rule, and a running experiment — can now change the same parameter, and nothing forces a clear line between them.

That ambiguity is where experiments quietly go wrong. If a config parameter is being personalised for one segment while an A/B test is also targeting it, a variant’s read of “uplift” might really be picking up someone else’s rollout. And because fetches are now the metered unit, a client that fetches Remote Config more aggressively than expected doesn’t just cost more — it can also destabilise variant assignment if caching isn’t handled consistently across app restarts.

Data Points to Track

  • Experiment ID tagged on every fetch, so it’s clear which config values came from an active A/B test versus a personalisation rule or a plain rollout
  • Variant assignment stability across app restarts and reinstalls, confirming a user stays in the same variant for the life of the experiment
  • Fetch call volume per user against the new free daily tier, to catch clients that are fetching far more often than the experiment design assumes
  • Exposure event firing rate — whether the SDK actually logs exposure every time a variant is applied, not just on first fetch
  • Parameter overlap conflicts where a personalisation rule and an A/B test target the same key for the same segment
  • Sample size and statistical significance per variant, tracked over the experiment’s run rather than checked once at launch

Setup Steps

  1. Audit existing Remote Config parameters to identify any already touched by more than one of A/B Testing, Personalization, or Rollouts.
  2. Tag every experiment-linked parameter explicitly so downstream analytics can separate experiment effects from personalisation or rollout effects.
  3. Instrument exposure logging to fire on every variant application, not just the first fetch after assignment, so exposure counts match actual usage.
  4. Monitor fetch volume against the free tier for your active user base, and set an alert before usage-based charges kick in unexpectedly.
  5. Set a fixed review cadence for sample size and significance per experiment, rather than calling a winner the first time a variant looks ahead.

Actionable Insights

The parameter overlap check is the one to run first: if it turns up any key targeted by both an experiment and a personalisation rule, every uplift number for that experiment is suspect until the overlap is resolved. Once experiment tagging and exposure logging are in place, the fetch-volume-versus-tier number tells you whether an experiment’s own design — too-frequent fetching, poor caching — is quietly pushing you into the metered pricing band, which is worth fixing before it becomes a recurring cost rather than a one-off surprise.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert