Docs

Smart Glasses Apps: What Counts as an Event

Smart glasses SDKs like Meta's Wearables toolkit break the tap/scroll event model — here's what to track when there's no screen to tap.

Engagement

Meta’s Wearables Device Access Toolkit moved from developer preview toward general availability this year, giving third-party developers real SDK access to build companion experiences for Ray-Ban Display glasses, alongside Google’s parallel Android XR push with hardware partners. AI smart-glasses shipments grew roughly 89% year-on-year in the most recent quarter. That’s a new tracking surface arriving fast, and it doesn’t fit the event taxonomy most mobile teams already have.

The problem is structural, not cosmetic. Screen views, tap coordinates, and scroll depth all assume a rectangle of pixels the user is looking at and touching. A glanceable heads-up display, a voice command, or a camera-triggered action doesn’t map onto any of those primitives cleanly. Teams that try to force wearable interactions into their existing mobile event schema end up with data that technically logs something on every interaction but tells you almost nothing about whether the interaction actually worked, because the schema was built for a different interaction model.

Get this wrong and you end up flying blind on the interactions that matter most for a wearable: did the voice command get understood, did the glanceable card get read or dismissed unseen, did the camera-triggered action fire when the user actually wanted it to. None of that is visible in a screen-view count.

Data Points to Track

  • Interaction modality, tagging each event as voice, glance, gesture, or camera-triggered rather than assuming a single tap-based event type
  • Voice command outcome, logging recognised vs. misrecognised vs. no-response separately from whether the resulting action succeeded
  • Glance duration and dismissal method, distinguishing a card that was read and consciously dismissed from one that simply timed out unseen
  • Camera-trigger confirmation, recording whether a camera-initiated action was confirmed, cancelled, or auto-expired
  • Session boundary definition, since a wearable session may have no clear “app open” moment the way a mobile app does

Setup Steps

  1. Define modality-specific event types before instrumenting anything — don’t reuse a mobile tap event with the coordinates blanked out.
  2. Instrument voice recognition confidence and outcome as a distinct field, separate from downstream action success, so misrecognition and action failure don’t get conflated.
  3. Track glance-level engagement (viewed, read, dismissed, timed-out) as its own funnel rather than folding it into a generic “impression” count.
  4. Establish a session-start and session-end signal appropriate to the device — a worn/removed state or an explicit wake event, not an app-foreground proxy.
  5. Version your wearable event schema separately from your mobile schema from day one, since the two will diverge quickly as the platform matures.

Actionable Insights

Once modality-aware events are in place, the data separates two failure modes that look identical in a generic interaction count: the assistant heard you correctly but took the wrong action, versus the assistant never understood you at all. That split determines whether the fix is in voice recognition tuning or in downstream action logic. Glance-level dismissal data does something similar for passive engagement — a card dismissed quickly after being read is a different signal from one that timed out unseen, and conflating them into a single “impression” metric hides which glanceable content is actually landing with users on a device where attention is measured in seconds, not minutes.

Expert help

Need help tracking this in your app?

Our team sets up analytics pipelines for mobile and web teams every day. Talk to us and get your first events flowing in under an hour.

Talk to an expert