Docs

Tracking Your App Inside ChatGPT (Since OpenAI Won't)

OpenAI's Apps SDK has no analytics dashboard yet — what to self-instrument to see how your ChatGPT app is actually discovered, opened, and used.

Analytics

OpenAI’s Apps SDK, built on the Model Context Protocol, lets developers ship apps that run inside ChatGPT’s chat interface — fetching data, rendering UI, and taking actions without the user leaving the conversation. With over 800 million ChatGPT users as the addressable audience, teams are moving fast to get an app into the directory. What they’re finding once it ships is a gap App Store and Play Console developers haven’t dealt with in over a decade: there’s no first-party analytics dashboard. OpenAI’s developer community has been asking for basic numbers — impressions, traffic, how an app gets surfaced — since apps started rolling out, and that data still isn’t exposed anywhere.

That leaves teams flying blind on the question that matters most: whether the app is being found and used, or quietly ignored inside a conversation the developer never sees. A traditional app has an install event, a session start, a store listing view, all instrumented before you write a line of tracking code. A ChatGPT app has none of that by default. If a model decides not to invoke it in a relevant conversation, or a user dismisses its output, nothing surfaces that anywhere unless the app’s own backend logs it. Until OpenAI ships first-party analytics, self-instrumentation is the only way to know the product is working at all.

Data Points to Track

  • Tool invocation events, logged server-side every time the model actually calls one of the app’s exposed tools, since this is the closest equivalent to a session start and the only reliable signal that the app was surfaced rather than ignored
  • Invocation-to-completion rate, tracking how often a triggered tool call returns a result the user acts on versus one that gets cut off, ignored, or superseded by the model choosing a different response path
  • Widget interaction events, for any in-conversation UI the app renders, since a rendered card that nobody taps or scrolls past is functionally the same as not rendering at all
  • Conversation context at invocation, captured only as broad category or intent rather than raw message content, to understand which kinds of user requests actually trigger the app versus which similar requests don’t
  • Downstream action completion, whatever “done” means for the specific app — a booking confirmed, a query answered, a purchase completed — tied back to the invocation that started it, since ChatGPT’s own interface won’t report this for you

Setup Steps

  1. Log every tool call server-side, not client-side, since the app’s own backend is the only place with guaranteed visibility into what the model actually invoked and when — there’s no equivalent of a client SDK auto-firing session events here.
  2. Assign a session or conversation identifier at first invocation and thread it through every subsequent tool call in the same conversation, so a multi-step interaction can be reconstructed as one journey rather than a set of disconnected calls.
  3. Instrument widget render and interaction separately from tool invocation, using whatever callback or action hooks the Apps SDK exposes for in-conversation UI, so “the app ran” and “the user engaged with what it showed” are distinguishable in the data.
  4. Pipe logs into your own analytics warehouse immediately, treating the Apps SDK the same way you’d treat any headless integration with no vendor dashboard — because right now, there isn’t one to fall back on.
  5. Set up a weekly manual audit of invocation volume and completion rate while the platform’s own reporting is still absent, since without it, a real drop in usage and a temporary platform issue look identical unless someone is checking the numbers by hand.

Actionable Insights

A high invocation rate paired with a low completion rate points at the app itself — either the tool’s output isn’t matching what the model expected to show, or the widget UI isn’t clear enough to prompt action, both fixable without waiting on OpenAI. A low invocation rate despite the app being live in the directory is a discovery problem, not a product problem, and worth cross-checking against the prompts that should plausibly trigger the app but aren’t. And because there’s no official baseline yet, the first few weeks of self-instrumented data are themselves the baseline — a calibration period, not a comparison against an industry benchmark that doesn’t exist for this surface yet.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert