Statsig joined Amplitude earlier this year, and the integration work is now landing: Amplitude’s behavioural data can feed directly into Statsig’s feature flags and experiments, and Statsig’s exposure events can flow back into Amplitude’s charts. The pitch is a single system where a product team ships a flag, measures its effect on retention or activation in the same tool that already holds the rest of the funnel, and doesn’t have to stitch together two separate event schemas to answer “did this experiment actually work.”
The risk sits exactly at that seam. Exposure events (a user was assigned to a variant) and outcome events (a user did the thing you’re measuring) come from two systems that didn’t share an event taxonomy until now, and a naive integration will double-fire exposure logging, mismatch user identity between Statsig’s device ID and Amplitude’s identified user ID, or attribute an outcome to the wrong variant because the exposure event arrived after the behavioural event it’s meant to explain. Teams that skip validating the join end up trusting an experiment result that’s actually measuring assignment noise, not product impact — and a false positive here ships a change to everyone based on a broken read of a small test.
Data Points to Track
- Exposure event timing relative to the first behavioural event it’s meant to precede, to catch out-of-order delivery
- Identity resolution rate between Statsig’s assignment ID and Amplitude’s user ID, especially for anonymous-to-identified transitions mid-experiment
- Duplicate exposure count per user per experiment, since a retry or re-render can log the same assignment twice
- Variant balance drift — whether the actual traffic split matches the configured split, which flags a targeting or SDK bug before it corrupts results
- Cross-tool event volume parity between what Statsig logs as exposures and what Amplitude receives, to catch dropped events at the pipe
- Guardrail metric movement (crash rate, load time, error rate) alongside the primary experiment metric, so a “win” isn’t hiding a regression elsewhere
Setup Steps
- Map identity resolution explicitly between Statsig’s assignment identifier and Amplitude’s user ID before running any joint experiment, rather than trusting default matching.
- Instrument exposure logging once, at the point of variant assignment, and confirm no duplicate fires from retries, re-renders, or SDK re-initialisation.
- Run a synthetic A/A test (two identical variants) through the integration first to confirm the measured lift is statistically indistinguishable from zero before trusting real A/B results.
- Set a variant balance alert that fires if actual traffic split deviates from configured split by more than a few percentage points.
- Pair every experiment dashboard with its guardrail metrics, not just the primary metric, so a rollout decision accounts for the full picture.
Actionable Insights
Once exposure and outcome events are validated against each other, the number worth watching is the gap between Statsig’s reported exposure count and the matching identified-user count in Amplitude. A small, stable gap means the identity bridge is solid and experiment results can be trusted for rollout decisions; a growing or inconsistent gap means variant assignments are leaking or being misattributed, and any lift you’re seeing in the combined dashboard needs to be re-verified against raw event counts before it justifies shipping a flag to 100% of users.
Related Resources
Need help tracking this in your app?
Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.
Talk to an expert