Instrumenting the Vercel AI SDK With OpenTelemetry
Catch silent AI failures by tracing agent reasoning with structured observability.

Instrumenting an AI agent with OpenTelemetry starts from a specific failure: a request can return a 200 OK, the latency chart can stay flat, and three steps deep inside the agent's reasoning a tool can fire with the wrong argument. No traditional dashboard will catch that. The gap is architectural, a consequence of monitoring tools built for deterministic systems being asked to judge something they have no mechanism to judge. HTTP-level monitoring was designed for deterministic systems, where a success code reliably means the system did what it was asked to do. Agents break that assumption, because they can return a clean 200 and a fabricated answer in the same response. Request-level application performance monitoring reads status codes and response times; it has no column for "the answer was fabricated" or "the tool call used the wrong identifier." Vercel's AI observability guide names this directly as a key takeaway: traditional APM misses the failure mode that matters most in LLM applications, where a confidently wrong answer and a healthy-looking dashboard arrive together. Catching it takes a different kind of signal than status codes and latency, one that follows the reasoning path an agent takes, not just the request it received.
What the Vercel AI SDK emits over OpenTelemetry
The Vercel AI SDK, with telemetry enabled, produces a structured trace for every call: the full operation, each provider call within it, every tool execution, and every agent step, all captured as part of one coherent record. That structure turns a trace into something a debugging engineer can actually use to follow what the agent did.
For each trace, the SDK captures the full operation duration and per-provider-call timing, so you can trace a slow agent run back to a specific model call. It records the provider name, the requested model, and the actual response model, and that matters when a fallback or routing decision silently substitutes one model for another. Token counts, input, output, and cached, are recorded per step, along with the temperature and generation settings that produced that step's output. Finish reasons and response IDs get captured too, so each generation has a traceable identity. The detail that matters most for catching silent misbehavior is the tool name, arguments, and result recorded for every tool execution: this is what makes it possible to see that a tool fired with the wrong argument. HTTP status and error messages are captured on failures, and prompt and completion text is captured by default, giving a reviewer the actual words the model saw and produced, not just metadata about them.
These spans aren't a flat list. They follow a hierarchy defined by the OpenTelemetry GenAI Semantic Conventions: an invoke_agent span at the top, with chat spans beneath it for each language model call, and execute_tool spans for each tool invocation. If a backend understands gen_ai.* attributes, it can use that hierarchy to reconstruct the full reasoning chain an agent followed, not just a list of isolated calls stripped of their order. Namespace matters here: OpenTelemetry's GenAI conventions use a gen_ai. attribute prefix, while the OpenInference convention uses llm., and which one an exporter emits determines which backends can parse the resulting spans. When an agent runs multiple steps, you can attach custom span data at each step boundary with onStepStart and onStepEnd callbacks, so the default trace grows with whatever context you need to debug.
How AI SDK 7 changed the instrumentation architecture
AI SDK 7 moved telemetry registration out of the call site and into application startup, and that relocation changes how the rest of this setup has to be built. The new package is @ai-sdk/otel, which exports an OpenTelemetry class, and the canonical entry point is registerTelemetry, imported from the 'ai' package itself; you call it once when the application starts.
The practical consequence is strict ordering: telemetry has to be registered before any AI SDK code runs, not configured per request. Tracing setup needs to be imported before any AI SDK calls happen anywhere in the application, because the SDK will otherwise run without the hooks that registerTelemetry installs. Custom tracers follow the same shift. Instead of passing them as an argument on each individual call, you supply them once, to the OpenTelemetry constructor, and they apply globally from that point forward.
Some integrations still document the older pattern, an experimental_telemetry: { isEnabled: true } flag set per call, which reflects the pre-v7 approach. You can still use it where backward compatibility matters, but you shouldn't build new setups on it. AI SDK 7 also requires ESM, a build-system constraint distinct from the OTel wiring itself but one that directly affects where the instrumentation file can live and how the runtime loads it. Together, these two changes, registration at startup instead of at the call site, and a mandatory module format, are what the setup steps below are built around.
Installing the required packages and understanding what each one does
The instrumentation stack is built from three distinct layers: the AI SDK's own OpenTelemetry integration, the OpenTelemetry Node SDK that manages the spans once they exist, and an OTLP exporter that ships those spans somewhere. Skipping or confusing any one of the three produces a setup that starts up cleanly, runs without error, and exports nothing at all, which is a harder failure to diagnose than an outright crash.
For a Node.js deployment, the minimal install is:
- @ai-sdk/otel, the AI SDK's OpenTelemetry integration, which provides the OpenTelemetry and LegacyOpenTelemetry classes along with the registerTelemetry function
- @opentelemetry/sdk-node, the OpenTelemetry Node.js SDK, which manages span processors and the SDK's lifecycle from startup to shutdown
- @opentelemetry/exporter-trace-otlp-grpc, which ships spans over gRPC to an OTLP-compatible backend (@opentelemetry/exporter-trace-otlp-http is the alternative for HTTP/JSON transport, and @opentelemetry/exporter-trace-otlp-proto for HTTP/protobuf)
For a Vercel-hosted application, the package list is different. Vercel's instrumentation documentation calls for @vercel/otel and @opentelemetry/api. The @vercel/otel package wraps the OpenTelemetry SDK, for both the Node.js and Edge runtimes, with platform defaults already applied, so the registerOTel call handles environment-specific wiring that would otherwise need to be configured by hand.
The choice between gRPC and HTTP exporters isn't a style preference. It has to match whatever protocol the receiving backend expects. gRPC is the default assumed in a standard Node.js setup, but some backends require HTTP/JSON transport specifically. The exporter package and the OTEL_EXPORTER_OTLP_PROTOCOL setting both need to agree with the backend's expectations, because the two transports aren't interchangeable at runtime.
Wiring up the instrumentation file: startup order, registration, and the ESM constraint
The single most common mistake in this setup is placing OpenTelemetry initialization after the AI SDK has already been imported elsewhere in the application. The SDK registers its spans at module load time, so any tracing configuration that arrives after that point is silently ignored for every call already initialized under the old configuration. There's no error thrown. The traces fail to appear where they're expected.
The fix is a dedicated instrumentation file, conventionally named tracing.ts or tracing.mjs, and you import it before anything else runs in the application's entry point. The sequence inside that file matters and should be thought of as four ordered steps. First, import NodeSDK from @opentelemetry/sdk-node. Second, import the exporter, typically OTLPTraceExporter, from whichever exporter package matches the target backend's protocol. Third, import registerTelemetry from the 'ai' package and OpenTelemetry from @ai-sdk/otel. Fourth, instantiate NodeSDK with the exporter configured, call sdk.start(), and only then call registerTelemetry(new OpenTelemetry()). If that order is reversed, if registerTelemetry runs before sdk.start(), or if either runs after the rest of the application has already begun issuing AI SDK calls, spans either go unregistered or get dropped before the exporter is listening for them.
Vercel-hosted Next.js applications use a different entry point. Rather than a standalone tracing file, the convention is an instrumentation.ts file at the project root (or inside src/ if the project uses that layout), exporting a register() function that calls registerOTel({ serviceName: 'your-project-name' }). Next.js calls this function automatically once, when a new server instance starts, and the framework guarantees it finishes before the server starts handling requests, so you don't have to enforce the manual ordering the Node.js path demands.
The ESM constraint applies regardless of which entry point is used. AI SDK 7 requires package.json to declare "type": "module", or requires individual files to use the.mjs extension. CommonJS's require() isn't supported, so if you try to use it, it either fails silently or throws at import time, depending on how the surrounding build tooling handles the mismatch.
The serviceName field passed into the SDK constructor or into registerOTel isn't a cosmetic label. It becomes a resource attribute attached to every span the application emits, and it's the field most backends use to group and filter traces by application. If you leave it unset, or set it inconsistently across services, multi-service tracing gets far harder to read later.
Configuring the OTLP exporter and environment variables
Once the SDK is wired up and registering spans in the right order, the exporter configuration governs whether those spans actually reach a backend. Three environment variables govern this, and all three have to match the receiving system exactly, because a mismatch in any one of them produces a setup that runs cleanly, emits spans locally, and delivers nothing.
The full URL of the backend's OTLP receiver, the address spans are sent to, is set by OTEL_EXPORTER_OTLP_ENDPOINT. OTEL_EXPORTER_OTLP_HEADERS carries the authentication credentials the backend expects, and the required format varies by backend: OpenObserve expects an Authorization header set to Basic followed by a base64-encoded token, while Humanloop's integration expects an X-API-KEY header set to the raw key. OTEL_EXPORTER_OTLP_PROTOCOL specifies grpc or http/json depending on what the backend accepts; Humanloop's integration specifies http/json explicitly, and sending gRPC-formatted spans to an HTTP/JSON receiver, or the reverse, produces no visible error on the sending side even as every span silently fails to land.
If you set these three values through environment variables rather than hardcoding them into the instrumentation file, that file stays portable across environments. The same tracing.ts can send spans to a local collector during development and to a production backend once deployed, with no code change, just a change in which environment variables are set at runtime.
For Vercel deployments specifically, context propagation across service boundaries is configured through instrumentationConfig.fetch. The propagateContextUrls list names the domains that should receive the traceparent header on outbound requests. The dontPropagateContextUrls list excludes third-party services that shouldn't receive trace context, and ignoreUrls suppresses tracing for internal tools that don't need to appear in the trace. Without this configuration, spans generated across service boundaries have no shared lineage: each service creates its own isolated spans with no connecting thread between them.
The traceparent header itself is the mechanism behind that connection. When an AI agent call triggers a downstream fetch to another service, the traceparent header carries the trace ID and span ID of the originating call, so the receiving service can attach its own spans to that same trace. Failing to propagate that header correctly turns a single agent run touching three different services into three disconnected traces instead of one complete record of what the agent actually did, undermining much of the purpose of instrumenting the agent.