← Back to Blog
DevOps & Dev Workflows

OpenTelemetry GenAI Tracing: A Practical Playbook for LLM and Agent Observability

OpenTelemetry GenAI tracing gives engineering teams a practical way to inspect LLM and agent workflows beyond ordinary service logs. This playbook explains what to instrument, what to redact, how to connect traces to evaluations, and how to start without creating noisy or risky telemetry.

Written by Hamza Diaz
October 4, 202610 min read21 views

Why OpenTelemetry GenAI tracing matters now

OpenTelemetry GenAI tracing gives engineering teams a way to inspect LLM and agent workflows without treating the model call as a black box. That sounds dry until something breaks. The API returns 200. The app does not crash. The user gets an answer. Still, the answer may be thin, too expensive for the task, built from the wrong retrieved context, or shaped by a tool call that nobody noticed in review.

That is the awkward part of production AI work. Standard logs can prove a request happened. They rarely explain why an AI workflow behaved the way it did.

OpenTelemetry GenAI is not another monitoring product to buy. It is a set of semantic conventions for describing generative AI operations in a consistent format. The official OpenTelemetry guidance covers model requests, responses, events, metrics, exceptions, agent spans, and related telemetry. If teams describe providers, models, prompts, completions, token usage, tools, and agent steps with shared conventions, traces become easier to query, compare, and move between observability backends.

My view: most AI teams wait too long to standardize this. They add tracing after the first painful incident, when prompt versions, routing rules, and tool behavior are already scattered across logs, notebooks, and dashboard screenshots. The better time is before a workflow becomes important.

From application logs to model, tool, and agent traces

Most production teams already monitor the basics: HTTP latency, status codes, database calls, queue depth, dependency failures, and error rates. LLM systems fail in quieter ways. A generated answer can be poor because a prompt template changed, retrieval selected weak sources, the model was routed to a different provider, a tool returned partial data, streaming stopped early, a retry changed the final context, or validation allowed an answer that should have been blocked.

Tracing gives those steps a shape. Instead of one opaque operation named generate answer, a trace can show prompt assembly, retrieval, model invocation, tool selection, tool execution, validation, and response assembly as related spans. That matters when a prototype turns into a managed workflow. Debugging, incident review, privacy controls, and cost review need more than screenshots and ad hoc JSON logs.

What the GenAI semantic conventions add

The OpenTelemetry GenAI conventions give teams a shared vocabulary for model operations. A model call can sit inside the same distributed trace as the web request, retrieval service, database query, and downstream tool. At a high level, the conventions cover the GenAI system or provider, operation names, request and response models, token usage, prompts and completions as events when captured, exceptions, metrics, and agent-related spans.

The limit is just as important as the benefit. Tracing helps teams reconstruct what happened during a request. It does not prove the final answer was correct, safe, or useful. You still need evaluations, red teaming where appropriate, access controls, provider-specific logs, privacy review, and product judgment.

The observability gap in LLM applications and agents

Traditional application performance monitoring is strong on service health. It can show that an endpoint was slow, a dependency timed out, or a database query failed. For LLM applications, those signals are necessary but incomplete. A normal service trace often misses prompt construction, prompt template version, retrieved document identifiers, model choice, sampling settings, stop reason, tool routing, retry behavior, safety decisions, and evaluation outcomes.

Agents widen the gap. They may plan, call tools, inspect results, call the model again, retrieve memory, hand off to another workflow, run validations, and stop under a termination rule. Some steps can fail while the user still receives a final answer. Teams working on agentic systems should also read more broadly about AI agent reliability and oversight, because tracing shows the path an agent took while oversight decides whether that path was acceptable.

A practical first version of LLM observability usually records model and provider, operation name, request model, response model when available, request parameters, latency, token counts where available, stop reason, tool name, tool result status, retrieval source identifiers, error categories, safety flags, and evaluation outcomes. Content capture needs stricter policy. Prompts and completions can contain user data, private documents, secrets, contracts, source code, health information, financial data, or internal strategy. Many teams should start with metadata, IDs, hashes, and redacted snippets rather than full content.

What OpenTelemetry GenAI standardizes

The OpenTelemetry semantic conventions define a common language for telemetry. In the GenAI area, they describe attributes and events that make model operations recognizable across services and tools. A model invocation span can identify the GenAI provider, operation name, requested model, response model, request parameters, token usage, and response metadata. The exact attribute names should be checked against the current OpenTelemetry GenAI repository before implementation because the conventions are still changing.

Events are useful because a model request is not always well represented as static span attributes. A chat interaction may contain several messages. A streaming response may produce chunks over time. A tool call may be associated with an assistant message and a later tool result. Security still comes first. Events can carry sensitive text if content capture is enabled. Teams building document workflows may find the Optijara guide to document AI pipelines useful, since document systems often combine retrieval, extraction, generation, and privacy constraints.

OpenTelemetry Protocol, usually shortened to OTLP, gives teams a route to send telemetry to compatible observability backends. Backend support varies. Some tools display GenAI traces clearly. Others treat them as ordinary spans with attributes. Before calling this production-ready, verify queryability, trace visualization, retention behavior, and access control in the backend your team actually uses.

The Optijara GenAI Trace Map

The Optijara GenAI Trace Map is a decision-first framework for designing LLM and agent observability. It turns tracing from a data collection habit into a production design activity. The question is not what can we log. The better question is what decision should this trace help us make, which spans expose that decision, and which content must stay out of telemetry.

flowchart TD A[User request] --> B[Prompt assembly] B --> C[Retrieval or memory lookup] C --> D[Model call] D --> E{Tool needed?} E -->|Yes| F[Tool execution] F --> G[Tool output validation] G --> D E -->|No| H[Response validation] H --> I[Response assembly] I --> J[Optional eval signal] J --> K[Incident, product, and cost review]

Start with decisions, not dashboards. Useful trace questions include which model call was slow, which provider handled the request, which prompt template version was used, which documents were retrieved, which tool failed, whether retries changed the final answer, whether validation passed, and which workflow produced unusual usage.

Trace layerTypical spanUseful metadataAvoid by default
EntryUser requestTrace ID, route, user segment, request typeFull user message without policy
PromptPrompt assemblyTemplate ID, version, context IDs, redaction statusRaw confidential context
RetrievalSearch or rerankIndex name, document IDs, score bands, result countFull private documents
ModelGenAI callProvider, model, parameters, latency, token usageSecrets in prompts or completions
ToolTool executionTool name, status, duration, error categoryTool credentials or raw sensitive output
ValidationSafety or schema checkPass or fail, rule ID, fallback usedPrivate policy text if restricted
EvaluationQuality signalEval name, rubric version, score bandTreating eval as absolute truth

A compact machine-readable policy can help engineering and governance teams align before instrumentation spreads:

{
  "framework": "Optijara GenAI Trace Map",
  "trace_goal": "debug and evaluate one production AI workflow",
  "required_spans": ["entry", "prompt_assembly", "retrieval", "model_call", "tool_call", "validation", "evaluation"],
  "content_policy": {
    "forbidden": ["secrets", "credentials", "unredacted private documents"],
    "redacted": ["user text", "tool output"],
    "sampled": ["prompt events", "completion events"],
    "retained": ["metadata", "trace ids", "template versions", "document ids"]
  }
}

Implementation patterns, from one LLM call to agents

The simplest implementation traces one model request inside an existing application trace. The parent span is the user request or background job. The child span is the GenAI operation. It should capture the provider or system, requested model, operation name, relevant parameters, latency, status, token usage where available, response model where available, and exception details if the call fails.

Retrieval augmented generation needs more than a model span. A useful RAG trace usually includes query rewriting, embedding search, reranking, selected document IDs, context assembly, final generation, citation validation, and response assembly. Document IDs matter because they let teams audit which sources influenced the answer without placing full confidential documents into telemetry.

For agents, represent the workflow as a tree. The root span is the user task. Child spans can include planning, model calls, tool selection, each tool execution, tool output validation, reflection, memory reads, response validation, and termination. If the agent calls tools in parallel, each tool call should be visible as a sibling span with status and duration.

user_task: investigate_invoice_question
  prompt_assembly: support_agent_v4
  model_call: plan_next_step
  tool_call: search_orders status=ok
  tool_call: fetch_invoice status=timeout
  model_call: revise_plan_after_timeout
  tool_call: fetch_invoice status=ok retry=1
  validation: policy_and_schema_check status=passed
  response_assembly: final_answer

LangSmith documents OpenTelemetry tracing support, which is a useful example of ecosystem integration. Framework instrumentation can reduce manual span work, especially when teams already use a framework for chains, tools, or agents. Treat that instrumentation as a starting point, then inspect the exported traces. This pairs well with broader engineering hygiene such as using uv for Python project workflows, where reproducible tooling and consistent environments make observability rollout easier to maintain.

Adoption and evaluation plan

A safe rollout starts with one workflow and a test checklist. Verify trace continuity across services. Confirm that GenAI attribute names match the current OpenTelemetry documentation your team has chosen to follow. Check sampling behavior. Test prompt redaction. Inspect streaming responses. Confirm whether token usage is available from your provider and how it is represented. Make sure tool errors are captured as tool errors, not only generic exceptions. Query traces in the backend and confirm that on-call engineers can find what they need.

Test areaWhat to verifyWhy it matters
Trace continuityOne trace follows the request across app, retrieval, model, and toolsPrevents split diagnosis
Attribute namingNames align with chosen OpenTelemetry GenAI versionImproves portability and queries
RedactionSensitive prompts, documents, and tool outputs are filteredReduces privacy and security risk
SamplingHigh volume workflows do not overwhelm storageControls telemetry cost and noise
StreamingChunks, final response, and stop reason are represented consistentlyMakes streaming failures diagnosable
Token metadataToken fields are present where providers expose themSupports usage review with caveats
Backend queriesEngineers can search by trace ID, model, workflow, and error typeMakes telemetry usable during incidents

Avoid capturing full prompts by default. Avoid relying on a single dashboard as the source of truth. Avoid treating token counts as exact across every provider and workflow. Avoid instrumenting every helper function. Avoid ignoring retention and access policy. Avoid using traces as proof that an answer was good. OpenTelemetry conventions continue to evolve, so teams should pin instrumentation versions, document assumptions, and review convention changes before broad rollout.

Common mistakes teams make

The fastest way to create observability risk is to log prompts and completions first, then discuss policy later. Prompts can contain personal data, confidential documents, credentials, source code, private business plans, or regulated information. Define forbidden, redacted, sampled, and retained fields before instrumentation.

Model-only tracing is too narrow for agents. Many failures happen outside the model: retrieval returns weak context, a tool times out, validation is skipped, memory contains stale information, or response assembly drops important detail. If the trace stops at the model span, the team may blame the wrong layer.

Every attribute should answer a question. Who will use this field, during which workflow, and what decision will it support? If nobody can answer, the field is probably noise. Traces are evidence of process. Evaluations are evidence of output quality, and even evaluations need careful rubric design.

How to choose your first tracing project

The best first project is not always the most complex AI feature. Choose a workflow where observability will change decisions soon.

CriterionLow priority signalHigh priority signalFirst project guidance
Business importanceInternal experimentUser facing or operationally important workflowPrefer high impact but bounded scope
Debugging painFailures are rare or obviousFailures are frequent, ambiguous, or costly to investigateStrong candidate
Privacy sensitivityMostly public or synthetic dataSensitive user or document dataStart only with strict redaction policy
Workflow complexitySingle model callRetrieval, tools, validation, retries, or agentsTrace map adds more value
OpenTelemetry maturityNo existing tracingExisting distributed traces and OTLP pipelineEasier rollout
Backend readinessNo query or access modelQueryable traces with access controlsSafer production adoption

A practical first milestone is one end-to-end request with a readable trace tree. The trace should connect to application logs through trace ID. It should show prompt assembly, retrieval if present, model call, tool outcomes if present, response validation, and one evaluation or review signal. Redaction should be verified before any wider rollout.

If your team is moving from prototypes to production AI workflows, Optijara can help map the traces, evals, and governance checks before instrumentation becomes noisy or risky. The aim should be vendor-neutral observability that helps teams debug, evaluate, and operate AI systems with clear boundaries.

OpenTelemetry GenAI tracing is most useful when teams standardize before scaling, trace decisions rather than every possible call, protect sensitive content, connect traces to evaluations, and keep the conventions under review as the ecosystem matures.

Key Takeaways

  • 1OpenTelemetry GenAI tracing standardizes how teams describe model calls, prompts, completions, tools, agents, token usage, and provider metadata.
  • 2Traces are most valuable when designed around debugging, evaluation, cost review, governance, and incident analysis decisions.
  • 3Agent observability needs parent child spans for planning, retrieval, tool calls, validation, retries, and response assembly, not only model calls.
  • 4Prompt and completion capture should be governed by policy, redaction, sampling, retention limits, and access controls.
  • 5OpenTelemetry conventions improve portability, but teams still need to verify backend support, queryability, and visualization quality.

Conclusion

OpenTelemetry GenAI tracing gives production teams a shared language for inspecting LLM and agent workflows, but the standard only helps when it is applied with discipline. Start with the decisions your traces must support, define span boundaries before writing instrumentation, protect sensitive content by default, and connect traces to evaluations, incidents, and rollout reviews. That keeps observability practical instead of noisy.

Frequently Asked Questions

What is OpenTelemetry GenAI?

OpenTelemetry GenAI is a set of semantic conventions and telemetry patterns for tracing generative AI operations such as model calls, prompts, completions, token usage, tool calls, exceptions, metrics, and agent workflow spans.

How is LLM observability different from traditional application monitoring?

Traditional monitoring tracks service health, latency, and errors. LLM observability also needs visibility into prompt assembly, retrieval context, model parameters, tool calls, token usage, validation outcomes, and quality signals.

Should teams store prompts and completions in traces?

Not by default. Teams should define policy first, then use redaction, sampling, retention limits, and access controls. Metadata, identifiers, hashes, and redacted snippets are often safer than full content.

Can OpenTelemetry trace multi step AI agents?

Yes. Teams can model agent workflows as parent child spans for planning, prompt assembly, model calls, retrieval, tool execution, validation, retries, memory reads, and response assembly.

Does OpenTelemetry replace evaluation tools for LLM applications?

No. Traces explain what happened during a request. Evaluations help judge whether the output was correct, safe, useful, and aligned with the task. Production teams usually need both.

Sources

Share this article

Hamza Diaz

Written by

Hamza Diaz

Hamza Diaz is the founder of Optijara, where he builds practical AI agents, automation systems, and Copilot workflows for service businesses. He writes about AI operations, agent strategy, and real-world implementation for teams that want usable systems instead of hype.