Docs

Sentry Agent Tracing: Tracking AI Conversations

What to track in Sentry's Agent Tracing and Conversations view — transcript, timeline and span data for debugging AI agents in production.

Performance

Sentry’s redesigned Conversations view — built on its Agent Tracing product and expanded this month with automated triage via Claude Routines — turns a wall of raw AI spans into a chat-style replay of exactly what a user and an agent said to each other, with every tool call and model generation clickable underneath. That matters because AI agent failures rarely look like normal application errors. An agent doesn’t crash; it just quietly calls the wrong tool, hallucinates a parameter, loops on a failed step, or gives a confident but wrong answer. Without conversation-level tracing, none of that shows up in a stack trace, and support tickets become the only signal that something went wrong — usually days after the fact and with no way to reproduce it.

The gap teams hit first is instrumentation scope. It’s easy to trace the model call and miss everything around it: which tool was selected, what arguments were passed, how long a retrieval step took, and whether a downstream tool call actually succeeded. A transcript that shows only “user asked X, model replied Y” without the tool layer looks complete but tells you nothing about why an agent took five seconds too long or returned the wrong data. Getting agent observability right means capturing the full run — every span, not just the visible conversation — so a bad outcome can be traced back to the exact step that caused it.

Data Points to Track

  • Agent run ID, grouping every span (model calls, tool calls, retries) that belongs to a single end-to-end conversation
  • Tool call name, input arguments and output, logged for every tool invocation, not just the ones that error
  • Model generation metadata — model name, prompt and completion token counts, latency, and temperature/parameters used
  • Turn-level outcome, marking whether each conversational turn resolved successfully, was corrected by the user, or was abandoned
  • Error and retry events within a run, including which step triggered a retry and how many attempts it took
  • Session-to-conversation linkage, tying an agent run back to the originating user session so conversation replay lines up with normal product analytics

Setup Steps

  1. Instrument the agent framework, not just the LLM SDK call — wrap tool execution and retrieval steps in spans, not only the final model response.
  2. Tag every span with a shared run ID so Sentry (or your tracing backend) can group tool calls, retries and model generations into one conversation.
  3. Capture structured tool inputs and outputs, not just a success/failure flag, so a bad result can be diagnosed without reproducing the run.
  4. Mark turn-level resolution state explicitly in your application code — don’t rely on inferring success from the absence of an error.
  5. Route agent errors through the same alerting pipeline as your other production errors, rather than a separate, easily-ignored AI-specific channel.

Actionable Insights

Once conversation-level tracing is in place, the transcript and timeline views answer questions that error logs alone can’t: which tool calls are slow enough to be the actual latency bottleneck, which failure patterns cluster around a specific tool rather than the model itself, and how often users have to correct or retry before an agent gets to a useful answer. That last signal — correction and retry rate per conversation — is usually the clearest proxy for agent quality you’ll get, and it’s only visible once tool calls and turns are tracked as first-class data rather than folded into a single opaque “AI response” event.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert