OpenTelemetry Collector Configuration for LLM Observability Honeycomb published a complete, annotated OpenTelemetry Collector configuration for LLM observability that receives traces over OTLP, normalizes OpenInference and OpenLLMetry schemas onto GenAI semantic conventions, redacts sensitive prompt and completion content, manages span volume without sampling away conversations, and exports to Honeycomb. The configuration places redaction, enrichment, and sampling in a single Collector pipeline between the instrumented application and the backend rather than in each application, and includes validation steps and a troubleshooting table. OpenTelemetry Collector Configuration for LLM Observability A complete, annotated OpenTelemetry Collector configuration for LLM observability: receiving traces over OTLP, normalizing OpenInference/OpenLLMetry schemas onto GenAI semantic conventions, redacting sensitive prompt and completion content, managing volume without sampling away conversations, exporting to Honeycomb, plus validation steps and a troubleshooting table. By: Mike Goldsmith https://www.honeycomb.io/author/mike Your LLM application emits telemetry unlike anything else in your stack. Model calls, tool invocations, retrieval steps, and token usage arrive as spans whose attributes carry entire prompts and completions. That data is bulky, it's full of user content you may not be allowed to export, and depending on which instrumentation each service uses, the same fact can arrive under different attribute names. The OpenTelemetry OTel Collector is the control layer that solves all three problems in one place, sitting between your instrumented application and Honeycomb. Instead of embedding redaction, enrichment, and sampling logic in every application, you configure it once in a pipeline that every span flows through. In this post, I'll walk through a complete Collector configuration for LLM observability https://www.honeycomb.io/resources/getting-started/what-is-llm-observability , from receiving traces over OTLP, through normalizing, enriching, and protecting sensitive content before it leaves your network, to managing volume and exporting to Honeycomb, with validation and troubleshooting steps at the end. Learn more about Honeycomb Intelligence Connect with our experts today. What does the OpenTelemetry Collector do? The OpenTelemetry Collector is a vendor-neutral service that receives telemetry, processes it, and exports it. Receivers accept data in a given protocol, most commonly OTLP. Processors act on the data in a configured order, renaming attributes, redacting values, filtering spans, or sampling traces. Exporters send the result to one or more backends. Extensions add operational capabilities like health check endpoints, and the service section wires everything into pipelines. For LLM observability, the flow is simple. Your application's OpenTelemetry SDK exports spans over OTLP to the Collector, the processor chain shapes and protects the data, and the OTLP exporter delivers it to Honeycomb. Every configuration decision in this guide is about what happens between those two ends. For the full component reference, see how to send data with the OpenTelemetry Collector https://docs.honeycomb.io/send-data/opentelemetry/collector . Why use an OpenTelemetry Collector for LLM observability? The strongest reason is data protection. Prompts, completions, retrieval chunks, and tool arguments routinely contain user data, credentials, and internal context. The Collector redacts or drops that content inside your network boundary using rules you define, before anything reaches a third party, and the policy lives in one reviewable config rather than scattered across application code. The second reason is volume. Message payloads dwarf ordinary span attributes, and agent workflows fan out into many spans per user request. The Collector allows you to filter noise and sample intelligently so cost stays predictable without losing the traffic you care about. The third is consistency. Many tools now emit the OpenTelemetry GenAI semantic conventions natively over OTLP, but older schemas from OpenInference, OpenLLMetry, and homegrown wrappers are still common, and a fleet of services rarely agrees on one instrumentation vintage. Even native emitters drift across semantic convention versions as SDKs update. The Collector is where all of it converges on one queryable shape, so you can normalize GenAI telemetry with the OTel Collector https://www.honeycomb.io/blog/fast-ai-feedback-loops-honeycomb-opentelemetry instead of fixing every producer. Finally, the Collector routes. One pipeline can serve Honeycomb and a second destination without any backend-specific logic in your applications. Before you configure the Collector A short checklist before touching YAML: - An instrumented application. Official OpenTelemetry GenAI instrumentation now covers direct client libraries anthropic, openai, google-genai and agent frameworks langchain, openai-agents, smolagents in the opentelemetry-python-genai https://github.com/open-telemetry/opentelemetry-python-genai repository, so custom-built apps get the conventions natively by adding a package. OpenInference and OpenLLMetry instrumentation also works, normalized later in the pipeline. - A Collector distribution with the right components. The contrib distribution includes everything used here. The gen ai normalizer processor is alpha stability in contrib, and is not in the Honeycomb Collector distribution yet, so use contrib or a custom build if you need it. - A Honeycomb environment and API key , and service.name set on every service. - Privacy decisions made up front. Decide which content attributes may leave your environment at all. It's much easier to loosen a redaction policy than to claw back exported data. One caveat worth knowing: the GenAI semantic conventions are still evolving in their own repository https://github.com/open-telemetry/semantic-conventions-genai , so attribute names may change between SDK versions. The configuration below is built to absorb that. Configure the OpenTelemetry Collector for LLM telemetry Here is the complete configuration. Each stage is explained in the sections that follow. extensions: health check: receivers: otlp: protocols: grpc: endpoint: 0.0.0.0:4317 http: endpoint: 0.0.0.0:4318 processors: memory limiter: check interval: 1s limit percentage: 80 spike limit percentage: 20 Optional: only needed when some services emit non-OTel GenAI schemas such as OpenInference or OpenLLMetry. gen ai normalizer: sources: - name: openinference remove originals: true transform: trace statements: - context: resource statements: - set resource.attributes "deployment.environment.name" , "${env:DEPLOY ENV}" redaction: allow all keys: true Mask sensitive values wherever they appear, including inside message content, while leaving the rest of the message intact. blocked values: - "sk- A-Za-z0-9 {20,}" API keys - " 0-9 {3}- 0-9 {2}- 0-9 {4}" US social security numbers - "4 0-9 {12} ?: 0-9 {3} ?" Visa card numbers Strictest option: mask message content attributes entirely. Conversation views like Agent Timeline lose the messages, so only reach for this when policy demands it. blocked key patterns: - gen ai\.input\.messages - gen ai\.output\.messages filter: traces: span: - attributes "http.route" == "/health" exporters: otlp http: endpoint: https://api.honeycomb.io headers: x-honeycomb-team: ${env:HONEYCOMB API KEY} sending queue: enabled: true sizer: items queue size: 100000 batch: flush timeout: 200ms min size: 8192 service: extensions: health check pipelines: traces: receivers: otlp processors: memory limiter, gen ai normalizer, transform, redaction, filter exporters: otlp http The processor order matters: memory limiter goes first so the Collector protects itself before doing any work. gen ai normalizer runs before anything that matches on gen ai. attributes, so every later stage sees canonical names. Redaction runs after normalization for the same reason; its rules target one set of names instead of one per framework. Note there is no batch processor as batching and retry are handled by the exporter's sending queue , which is the current recommendation over the standalone processor. Secrets stay in environment variables, never in the config file. Receive LLM traces over OTLP The OTLP receiver accepts traces over gRPC on port 4317 and HTTP on port 4318. Most OpenTelemetry SDKs default to one of these, so pointing an application at the Collector is usually a single environment variable, for example OTEL EXPORTER OTLP ENDPOINT=http://collector:4318 . Binding to 0.0.0.0 makes the receiver reachable from other containers and pods, which is what you want for a gateway-style deployment. If the Collector runs as a sidecar or on the same host as the application, bind to localhost instead to avoid exposing the port. In Kubernetes, expose the receiver through a Service and point every workload's OTLP endpoint at it. Enrich and normalize LLM context If your services use the official GenAI instrumentation libraries, their spans arrive already using the conventions, and the job here is enrichment. The transform processor above stamps deployment.environment.name on every resource. The same pattern adds a team owner, a cost center, or a tenant identifier you're permitted to keep, so you can slice token usage by whatever dimension your organization budgets on. For services on older, non-OpenTelemetry schemas, the gen ai normalizer processor rewrites their attributes onto the conventions. It has built-in mapping tables for OpenInference and OpenLLMetry and accepts user-defined tables for homegrown schemas, with type coercion so string token counts become integers. A sample of what the OpenInference source does: OpenInference attribute GenAI semantic convention llm.model name gen ai.request.model llm.provider gen ai.provider.name llm.token count.prompt gen ai.usage.input tokens llm.token count.completion gen ai.usage.output tokens session.id gen ai.conversation.id With remove originals: true the old names are deleted after mapping, so dashboards only ever see one name. For shapes a mapping table can't express, such as deriving gen ai.operation.name from span name patterns, pair it with the transform processor. The normalizer also stamps the scope schema URL it targets, which pays off later. Because the GenAI conventions still evolve, attribute names will move under you. For example, gen ai.usage.prompt tokens becoming gen ai.usage.input tokens is a real past example. The schema processor uses those schema URLs to translate telemetry between convention versions, and I've written about managing semantic convention migrations with the Collector https://www.honeycomb.io/blog/managing-opentelemetry-semantic-convention-migrations-collector in depth. Protect sensitive LLM data before export This is the section I'd get right before anything ships. Prompts and completions are user content. Tool arguments can carry credentials. Retrieval chunks quote your internal documents. The Collector is the last point where you control that data, so I think the redaction policy belongs here, not in application code you'd have to audit service by service. The redaction processor supports three complementary controls. blocked values masks matching substrings wherever they appear, so a card number or air miles reference inside gen ai.input.messages is masked while the rest of the conversation stays readable, and the same patterns catch API keys that leak into any other attribute. blocked key patterns masks the values of matching attribute keys entirely, the strictest option shown commented out in the config above. And hash function including hmac-sha256 replaces matches with a stable hash instead of a fixed mask, useful for user identifiers where you need to correlate activity without exporting the raw value. How far to go is your choice, and the tradeoff is what Honeycomb can show you afterwards. Keeping message content with targeted value masking, as the config above does, preserves conversation debugging and Agent Timeline's message view while known-sensitive patterns never leave your network. Masking gen ai.input.messages and gen ai.output.messages entirely protects the most, but conversation-centric workflows lose the messages themselves, so you can see that an agent misbehaved without seeing what it said. If policy says an attribute must never leave the environment, delete it with the transform processor's delete key rather than masking it. Whatever you choose, it happens inside your network boundary, before export. Honeycomb's guide to removing sensitive information before export https://docs.honeycomb.io/send-data/opentelemetry/collector/handle-sensitive-information covers more best practices. Control LLM telemetry volume and reliability memory limiter protects the Collector from payload spikes, and the exporter's sending queue batches and retries so transient network failures don't drop data. The filter processor removes traffic with no analytical value, health checks in the example, and the same OTTL syntax can drop zero-token spans or verbose sub-steps you never query. For sampling, the usual goal is to keep the signal and sample the noise. For GenAI data, our recommendation today is not to sample it at all. Honeycomb's sampling guidance for GenAI data https://docs.honeycomb.io/manage-data-volume/sample/gen-ai explains why: a conversation loses its value when turns are missing, so completeness matters more than volume for this traffic. Filtering noise with the filter processor and trimming verbose attributes give you most of the volume control without breaking conversations, and you can filter unnecessary telemetry https://docs.honeycomb.io/manage-data-volume/filter/filter-processor beyond the health-check example above. If the surrounding non-LLM services need sampling, the adaptive tail sampling processor https://www.honeycomb.io/blog/bringing-most-advanced-sampling-opentelemetry-collector we contributed to the OpenTelemetry Collector handles those pipelines while your GenAI traffic flows through complete. Export LLM telemetry to Honeycomb The otlp http exporter sends to https://api.honeycomb.io use https://api.eu1.honeycomb.io for EU teams with your API key in the x-honeycomb-team header, read from an environment variable. TLS is on by default for this endpoint, so no extra transport config is needed. The sending queue settings shown earlier handle batching, retry, and backpressure at the export boundary. Check the Honeycomb OTLP documentation https://docs.honeycomb.io/send-data/opentelemetry/collector for current endpoint and header values, since these are the settings most likely to change over time. Validate your LLM traces in Honeycomb Validation here means confirming the configuration did what you intended, not just that data arrived. First, confirm normalization landed. Run a query grouped by gen ai.request.model . If every service appears under canonical model names regardless of which instrumentation it uses, the normalizer and your native instrumentation agree. A COUNT grouped by gen ai.operation.name should similarly show values such as chat and execute tool rather than framework-specific span kinds. Second, confirm redaction applied. Search for a pattern your policy blocks, an API key prefix or a test card number you deliberately sent through. Zero results means the policy held. Spot-check a span and verify its Gen AI tab renders the message with the masked pattern in place of the sensitive value, and the rest of the message intact. Third, confirm the numbers are usable. SUM gen ai.usage.input tokens and SUM gen ai.usage.output tokens grouped by model give you the token spend view, and opening any agent trace should show tool and model spans correctly parented under the workflow span, with deployment.environment.name present from the enrichment stage. Finally, open a conversation in Honeycomb's Agent Timeline. It's built directly on the GenAI conventions, so it doubles as a conformance check: if your conversation turns, agents, and tool calls render on the timeline, your pipeline is emitting the right shape. The attributes it relies on most are gen ai.conversation.id on every span of a run and gen ai.agent.name on each agent's spans, and the instrumenting AI agents guide https://docs.honeycomb.io/send-data/use-cases/agents covers what to send. Both are worth confirming survive your redaction and filter rules. Troubleshooting common OpenTelemetry Collector problems - No traces arriving: Check the app's OTLP endpoint and port 4317 gRPC / 4318 HTTP , receiver binding, and Collector logs for receiver startup errors. - Authentication errors: The HONEYCOMB API KEY env var isn't set in the Collector's environment, or the key belongs to the wrong environment. - Missing gen ai. attributes: The normalizer's source doesn't match the instrumentation. Inspect a raw span's attribute names and add a user-defined mapping table if needed. - Redaction not applying: Check processor order. Redaction must run after the normalizer, and rules written against pre-normalization attribute names silently match nothing. - Data being dropped: memory limiter under pressure or the sending queue overflowing. Check the Collector's own metrics and raise limits or queue size. - Unknown component on startup: Your distribution doesn't include that processor. Use the contrib or Honeycomb distribution, or build a custom one with the components you need. OTel Collector best practices for production LLM systems - Keep secrets in environment variables and the config in version control, and test processor rules against sample payloads before deploying. A redaction rule that almost matches won't be noticed until someone audits the exported data. - Monitor the Collector itself. Send its internal metrics to Honeycomb and alert on queue growth, dropped spans, and memory pressure. - Re-validate redaction on every instrumentation change. New SDK versions add new attributes, and a new content-bearing attribute is a leak until your policy covers it. - Size memory realistically for message payloads, which are far larger than typical span attributes, and review your pipeline when the GenAI semantic conventions publish a new version. For a worked production example of these practices in a different domain, see this OpenTelemetry Collector implementation example https://www.honeycomb.io/blog/monitoring-windows-servers-opentelemetry-collector . Start observing LLM workflows with Honeycomb Native GenAI instrumentation or the normalizer for older schemas gets everything onto one set of attribute names, redaction keeps user content inside your boundary, filtering keeps volume proportional to value, and the OTLP exporter delivers it. From there, Honeycomb's querying does the interesting work, from token spend by model to debugging a single agent conversation. Start sending OpenTelemetry data to Honeycomb https://www.honeycomb.io/platform/opentelemetry and see what your LLM workflows are actually doing. Want to try it with your own traffic? Sign up for a free Honeycomb account https://ui.honeycomb.io/signup and start querying, or book office hours https://www.honeycomb.io/devrel/observability-office-hours with one of our Developer Advocates if you want a hand getting set up.