APIs, integration & security — in depth
Agent TracingLong read

OpenTelemetry GenAI Semantic Conventions for Agent Spans

Standard schema replaces vendor chaos with shared agent telemetry structure.

Staff Writer · · 10 min read
Cover illustration for “OpenTelemetry GenAI Semantic Conventions for Agent Spans”
Agent Tracing · October 9, 2026 · 10 min read · 2,314 words

Before OpenTelemetry's GenAI semantic conventions existed, every AI observability product defined its own trace shape. That fragmentation turned instrumentation into a bet on a vendor instead of an investment in the application itself: switching backends meant rewriting instrumentation code across every service that touched a model. Agent telemetry is also structurally different from the HTTP and database telemetry that engineering teams already know how to monitor. Prompts and completions are large text blobs, tool-call parameters change shape on every single invocation, and multi-step reasoning just doesn't fit into a fixed schema the way a status code or a query duration does. The conventions that already existed for web requests and database calls simply have nothing to say about an LLM call.

The mismatch runs deeper than shape. AI systems run on a token-based cost model, so if you watch request rates, you learn almost nothing about what you're actually spending; a service can hold steady traffic while its bill triples because responses got longer. Latency tracks token count and model size rather than CPU load or I/O wait, so the instincts you built from a decade of infrastructure monitoring point in the wrong direction here. A single exception metric is equally unhelpful in an agent system: when a multi-step run fails, that metric cannot say whether the failure happened during planning, during a tool call, or during the model's own response generation. And because prompts frequently carry personally identifiable information, capturing that content can't default to the always-on logging that works for a database query; it needs an explicit opt-in.

A second cost compounds all of this. Agent requests commonly cross three or more frameworks or gateways on their way to completion, and if each one emits its own proprietary trace shape, the result is three incomplete, incompatible views of a single interaction. There's no end-to-end picture to debug from, only fragments that have to be reconciled by hand. That was the gap a dedicated specification from the OpenTelemetry community set out to close, built by a working group focused on this exact problem: a shared, vendor-neutral schema for the operations that make up an AI agent's work.

What the GenAI semantic conventions standardize

The GenAI semantic conventions standardize two things: which operations in an agent's execution deserve their own span, identified by the attribute gen_ai.operation.name, and the attribute names attached to those spans. Everything lives under the gen_ai.* namespace, which keeps these fields unambiguous even when they sit in the same trace as ordinary HTTP and database spans.

Five named operations cover most of what a team will see in practice. The chat operation represents a model call, and because models run on remote services, it carries a CLIENT span kind; this is the common LLM round-trip that most instrumentation starts with. The embeddings operation covers a vector embedding call and records vector dimension in gen_ai.embeddings.dimension.count. The execute_tool operation covers a single tool or function call, carrying the tool's name and arguments, and most agent failures surface during this call in practice. The invoke_agent operation covers one turn of an agent's reasoning and action; as of v1.41.0, it splits into a CLIENT span kind for remote agent APIs such as OpenAI Assistants or AWS Bedrock Agents, and an INTERNAL span kind for agents that run in-process inside a framework. The create_agent operation is emitted when an agent gets constructed.

A handful of request and usage attributes show up in nearly every trace, and are worth knowing by name. gen_ai.request.model names the model being called, while gen_ai.response.model carries the specific variant actually used, including fine-tuned versions. gen_ai.usage.input_tokens and gen_ai.usage.output_tokens give standard token accounting, turning cost aggregation into straightforward arithmetic across vendors. gen_ai.response.finish_reasons records why the model stopped generating.

Message content, meaning prompts and completions, is opt-in by design and lives in span events as span attributes. That placement matters operationally: events can be filtered or dropped at the Collector level without touching a single line of application code, which gives teams a privacy-by-design answer to the PII risk built into prompt content.

The nesting structure is what turns these individual spans into a coherent picture of an agent's work. At the top sits a session span, which groups everything that happens across one customer interaction and can span multiple turns. Inside that session, each invoke_agent span covers one turn of reasoning. Inside each of those, chat spans appear as children, carrying the model used, its parameters, and token usage, while execute_tool spans sit as siblings with their own timing and outcome. In a multi-agent system, a supervisor's invoke_agent span contains the invoke_agent spans of the sub-agents it triggered, each with its own model and tool descendants beneath it. Context propagation here is standard OpenTelemetry behavior, so the tree holds together even when a sub-agent runs in a completely separate service.

Two histogram metrics are close to mandatory for production monitoring: gen_ai.client.operation.duration, measuring latency in seconds, and gen_ai.client.token.usage, measuring token consumption broken down by input and output. Together, span structure and these two metrics give a team both the shape of an agent's execution and the numbers needed to track its cost and performance over time.

The spec's current stability status and the repository migration teams must know about

As of July 17, 2026, no GenAI-specific span, event, metric, or attribute is marked Stable in the dedicated repository. The entire gen_ai.* surface is in Development status. Shared core attributes referenced in the same tables, such as error.type and server.address, are Stable, which makes the contrast sharper: the scaffolding around GenAI telemetry has settled, but the GenAI-specific fields themselves have not.

The repository that holds this material also changed homes recently, and teams building instrumentation today need to know where the authoritative source now lives. The main open-telemetry/semantic-conventions repository's v1.42.0 release, on June 12, 2026, deprecated and moved all gen_ai.* content into a new, dedicated repository, open-telemetry/semantic-conventions-genai. The following release, v1.43.0 on July 3, ships none of the GenAI material at all, confirming the move is final. The old documentation page at opentelemetry.io now simply points readers to the new location.

The dedicated repository, though, has no releases or tags yet as of July 17, 2026, and the schema-URL section of its README still reads TODO. That leaves the last versioned cut of the GenAI conventions at main-repo v1.42.0. A team can follow a specific commit, track a framework's release cadence, or pin to a dated snapshot, but there's no versioned GenAI-conventions release at the new repository yet, and no finalized schema URL to pin instrumentation against.

The official documentation states that no public timeline exists for stabilization, and that attribute names and structures may still change. The spec does provide a release valve for this uncertainty in the form of the OTEL_SEMCONV_STABILITY_OPT_IN environment variable, which lets teams manage version transitions deliberately. None of this means the spec is unsafe to build on. The core concepts, which operations get spans, how those spans nest, where token usage lives, have held steady across the recent churn even as individual attribute names changed around them. Building on the GenAI conventions today is a reasonable bet for a team that pairs it with real versioning discipline, which is the subject the next two sections work through in detail.

Six releases that account for most of the attribute churn teams will encounter

Diagram: Six Releases That Broke Attribute Names. Visualizes: Show a vertical timeline of six specific spec releases that each introduced a breaking attribute change, so readers can see the sequence of churn at a glance.

Most of the attribute-name instability a team will run into in production traces traces back to six specific releases. Each one changed something that will break an existing query or a piece of instrumentation if a team doesn't handle it explicitly.

Release v1.27.0, from August 2024, renamed gen_ai.usage.prompt_tokens and gen_ai.usage.completion_tokens to gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. Any token dashboard covering data that spans this boundary needs dual-field queries or a migration step to avoid silently undercounting usage from before the rename. Release v1.37.0 replaced gen_ai.system with gen_ai.provider.name, and at the same time replaced per-message events with aggregated gen_ai.input.messages, gen_ai.output.messages, and gen_ai.system_instructions attributes. This is the most visible split between older framework telemetry and the current convention, and it exists because per-message events had been flooding multi-turn conversations with fine-grained events that were genuinely painful to query at scale. Release v1.38.0, from October 2025, added the gen_ai.evaluation.result event, carrying gen_ai.evaluation.name, gen_ai.evaluation.score.value, gen_ai.evaluation.score.label, gen_ai.evaluation.explanation, and gen_ai.response.id for correlation. That event is the bridge connecting tracing data to evaluation harnesses, a connection that matters more as teams start running automated quality checks against live agent traffic. Release v1.41.0 is where invoke_agent split into CLIENT and INTERNAL span kinds, so remote agent APIs separated from in-process framework agents, and this changes how a team filters and aggregates agent-level spans. Release v1.42.0 deprecated all GenAI material in the core repository and moved it to the dedicated repository described above, changing the authoritative source for the spec without a versioned release waiting at the new address. Underneath all five of these is a sixth source of churn: the dedicated repository has shipped no tagged release since the move, so any team tracking "current" behavior by following commits rather than a pinned version is exposed to changes that haven't been formalized into a release.

Knowing this sequence in order is what makes it possible to write queries that hold up across a dataset that spans these boundaries, rather than queries that quietly drop half the data the moment one of these renames took effect.

The specification and the telemetry a backend actually receives are two different things, and the distance between them is what decides whether a team's queries work on day one. Popular frameworks emit several different generations of GenAI attributes at once, often from a single install, because it depends on configuration and version.

Strands Agents defaults to frozen, v1.36-era compatibility behavior, and you can opt into the newer conventions only if you track issue #877. In direct tracing, legacy gen_ai.system=strands-agents fields appear with no accompanying gen_ai.provider.name, token usage is reported under both the old and new attribute generations with identical values, and span names already follow the current convention, invoke_agent, chat, execute_tool. That mixture isn't a bug; it's what the transition machinery is designed to produce while a framework straddles two spec generations. Vercel AI SDK 7 represents the newer generation more completely: OpenTelemetry support moved into a dedicated @ai-sdk/otel package, which emits gen_ai.provider.name and gen_ai.usage.input_tokens directly, while the older ai.* namespace remains available as a legacy path for teams that haven't migrated their dashboards yet. OpenAI Agents SDK takes a different route entirely: it ships built-in tracing on its own architecture and does not emit the OTel GenAI conventions natively, so getting that data into an OpenTelemetry pipeline requires contributed instrumentation or a custom processor layered on top. LangSmith sits in the middle of the spectrum: its documentation and examples mix attribute generations, but its OpenTelemetry ingestion accepts both, which makes it safer to connect to than to document from a single code sample. The practical rule that follows is to inspect the trace a system actually stores rather than assume its schema from one example in a guide.

The phrase "using the OpenTelemetry GenAI semantic conventions" doesn't point to one fixed telemetry schema in 2026. A team instrumenting a real agent application should expect a mix of attribute generations arriving from a single framework, and that mix will shift depending on package version and configuration choices made well before any trace reaches a backend. Instrumenting only the chat span captures the model call but misses the tool-call planning, the sub-agent handoffs, and the surrounding loop that often accounts for most of an agent's wall-clock time, producing a trace that looks fast while the actual user experience feels slow. You close both gaps, the generational mismatch and the span-coverage gap, with the kind of versioning discipline the next section lays out.

Building on a moving spec without breaking your dashboards

Because the spec sits in Development status and frameworks emit mixed attribute generations simultaneously, treating schema versioning as a first-class engineering concern is what keeps dashboards intact.

Pin every layer of the stack independently: the framework, the instrumentation library, the OTel SDK, and the exporter each carry their own assumptions about which convention generation they speak. A minor framework upgrade, the kind that normally ships without much scrutiny, can silently rename attributes underneath a dashboard that was working perfectly the day before. Treating these four layers as independently versioned, rather than assuming they move together, is what prevents that kind of silent breakage.

Any instrumentation a team ships or documents should come with a dated, tested-with table recording what was running when it was verified: the framework at its exact package version, the OTel SDK and exporter at their exact versions, the backend version receiving the data, and the value set for OTEL_SEMCONV_STABILITY_OPT_IN, all tied to a specific date of observation. That table turns a vague claim of compatibility into something another engineer can reproduce or rule out.

During transition periods, queries need to cover both attribute generations at once. If token data spans the v1.27.0 boundary, you need queries against both gen_ai.usage.prompt_tokens and gen_ai.usage.input_tokens, because data written before the rename only exists under the old name. If provider data spans the v1.37.0 boundary, you need the same treatment across gen_ai.system and gen_ai.provider.name. The safer default is to prefer the current field in new dashboards while tolerating the duplication of a legacy query running alongside it, and to hold off on deleting that legacy query until the migration of historical data is confirmed complete.

Finally, use OTEL_SEMCONV_STABILITY_OPT_IN deliberately rather than leaving it at whatever a framework defaults to. Framework defaults can lag the specification by one or more releases, and relying on them implicitly means a team's telemetry schema is decided by someone else's upgrade cadence. Setting that variable explicitly, and documenting the value alongside the tested-with table described above, is what keeps a dashboard's behavior predictable even while the underlying spec keeps moving.

Sources

  1. Gen AI
Filed underAgent Tracing

More in Agent Tracing