Docs

Amplitude's AI Plugin: Auditing Auto-Instrumentation

Amplitude's AI Plugin lets an AI agent auto-instrument events and build dashboards from natural language — unaudited, it can quietly mis-tag data.

Analytics

Amplitude’s AI Plugin turns a connected AI agent into an end-to-end instrumentation tool: it can add event tracking to a codebase, generate daily and weekly briefings, run session replay audits, and produce charts and dashboards from a natural-language request, all through a single install. It’s a genuine step up in speed — a team can go from “we should track this” to a working event and a dashboard in one conversation, without a person hand-writing the tracking plan or the SQL.

The risk that speed introduces is that instrumentation added by an agent, from a natural-language prompt, doesn’t carry the same scrutiny a human engineer applies when writing a tracking plan by hand. An ambiguous prompt can produce an event that’s named plausibly but scoped wrong — firing on the wrong trigger, double-counting a user action, or omitting a property the team assumed would be there because the agent inferred a different definition of the metric than the one in anyone’s head. Because the output looks like normal instrumentation and slots into existing dashboards, a mis-scoped AI-generated event can sit in production for weeks before anyone notices the underlying event volume or property set doesn’t match what the metric name implies.

Data Points to Track

  • Instrumentation source per event (AI Plugin-generated vs. human-authored), tagged in the tracking plan or event metadata, so agent-added events are identifiable without cross-referencing commit history
  • Prompt-to-event mapping, keeping a record of the natural-language request that produced each AI-instrumented event, so a later audit can check whether the resulting event actually matches the original intent
  • Event volume deviation from comparable hand-instrumented events, flagging AI-generated events whose fire rate looks unusually high or low relative to similar human-authored events for the same feature
  • Property completeness on AI-generated events, checked against the properties a human reviewer would expect for that event type, since an agent can omit a property it didn’t infer as necessary
  • Dashboard-to-source traceability, recording which auto-generated dashboards and charts were built directly from AI Plugin output versus reviewed and adjusted by a person

Setup Steps

  1. Tag every AI Plugin-generated event and dashboard at creation time with its source, so instrumentation provenance is queryable rather than reconstructed after the fact.
  2. Route new AI-generated events through a review step before they’re treated as production-ready — a lightweight check against the original prompt and an expected property list, not a full manual re-instrumentation.
  3. Set a volume-deviation alert comparing each new AI-generated event’s fire rate over its first week against the range for structurally similar existing events, to surface scoping errors early rather than at the next dashboard review.
  4. Keep the natural-language prompt alongside the tracking plan entry for every AI-instrumented event, so anyone auditing the event later has the original intent to check it against.
  5. Periodically re-run the session replay audit feature against a sample of AI-generated events specifically, rather than only using it for general QA, to catch drift between what the event claims to measure and what the replay shows actually happened.

Actionable Insights

An AI-generated event whose volume is consistently far outside the range of comparable human-authored events is either catching something genuinely different (worth understanding) or firing on the wrong trigger (worth fixing) — the volume-deviation alert doesn’t tell you which, but it tells you where to look first. If AI-generated dashboards are being adopted and shared without adjustment at a similar rate to human-built ones, that’s a sign the auto-instrumentation is holding up under real use; if AI-generated charts are consistently edited or rebuilt before anyone relies on them, that’s a signal the natural-language-to-metric translation needs tighter review before this becomes the default path for instrumentation. Either way, provenance tagging turns “is our data trustworthy” from a codebase-wide question into a targeted one — you can check the events an agent touched without re-auditing the events it didn’t.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert