A complete, annotated OpenTelemetry Collector configuration for LLM observability: receiving traces over OTLP, normalizing OpenInference/OpenLLMetry schemas onto GenAI semantic conventions, redacting sensitive prompt and completion content, managing volume without sampling away conversations, exporting to Honeycomb, plus validation steps and a troubleshooting table.
By: Mike Goldsmith
Your LLM application emits telemetry unlike anything else in your stack. Model calls, tool invocations, retrieval steps, and token usage arrive as spans whose attributes carry entire prompts and completions. That data is bulky, it's full of user content you may not be allowed to export, and depending on which instrumentation each service uses, the same fact can arrive under different attribute names.
The OpenTelemetry (OTel) Collector is the control layer that solves all three problems in one place, sitting between your instrumented application and Honeycomb. Instead of embedding redaction, enrichment, and sampling logic in every application, you configure it once in a pipeline that every span flows through. In this post, I'll walk through a complete Collector configuration for LLM observability, from receiving traces over OTLP, through normalizing, enriching, and protecting sensitive content before it leaves your network, to managing volume and exporting to Honeycomb, with validation and troubleshooting steps at the end.
Connect with our experts today.
What does the OpenTelemetry Collector do? #
The OpenTelemetry Collector is a vendor-neutral service that receives telemetry, processes it, and exports it. Receivers accept data in a given protocol, most commonly OTLP. Processors act on the data in a configured order, renaming attributes, redacting values, filtering spans, or sampling traces. Exporters send the result to one or more backends. Extensions add operational capabilities like health check endpoints, and the service section wires everything into pipelines.
For LLM observability, the flow is simple. Your application's OpenTelemetry SDK exports spans over OTLP to the Collector, the processor chain shapes and protects the data, and the OTLP exporter delivers it to Honeycomb. Every configuration decision in this guide is about what happens between those two ends. For the full component reference, see how to send data with the OpenTelemetry Collector.
Why use an OpenTelemetry Collector for LLM observability? #
The strongest reason is data protection. Prompts, completions, retrieval chunks, and tool arguments routinely contain user data, credentials, and internal context. The Collector redacts or drops that content inside your network boundary using rules you define, before anything reaches a third party, and the policy lives in one reviewable config rather than scattered across application code.
The second reason is volume. Message payloads dwarf ordinary span attributes, and agent workflows fan out into many spans per user request. The Collector allows you to filter noise and sample intelligently so cost stays predictable without losing the traffic you care about.
The third is consistency. Many tools now emit the OpenTelemetry GenAI semantic conventions natively over OTLP, but older schemas from OpenInference, OpenLLMetry, and homegrown wrappers are still common, and a fleet of services rarely agrees on one instrumentation vintage. Even native emitters drift across semantic convention versions as SDKs update. The Collector is where all of it converges on one queryable shape, so you can normalize GenAI telemetry with the OTel Collector instead of fixing every producer.
Finally, the Collector routes. One pipeline can serve Honeycomb and a second destination without any backend-specific logic in your applications.
Before you configure the Collector #
A short checklist before touching YAML:
- An instrumented application. Official OpenTelemetry GenAI instrumentation now covers direct client libraries (anthropic, openai, google-genai) and agent frameworks (langchain, openai-agents, smolagents) in theopentelemetry-python-genai repository, so custom-built apps get the conventions natively by adding a package. OpenInference and OpenLLMetry instrumentation also works, normalized later in the pipeline.
- A Collector distribution with the right components. The contrib distribution includes everything used here. The
gen_ai_normalizerprocessor is alpha stability in contrib, and is not in the Honeycomb Collector distribution yet, so use contrib or a custom build if you need it. - A Honeycomb environment and API key , and
service.nameset on every service. - Privacy decisions made up front. Decide which content attributes may leave your environment at all. It's much easier to loosen a redaction policy than to claw back exported data.
One caveat worth knowing: the GenAI semantic conventions are still evolving in their own repository, so attribute names may change between SDK versions. The configuration below is built to absorb that.
Configure the OpenTelemetry Collector for LLM telemetry #
Here is the complete configuration. Each stage is explained in the sections that follow.
extensions:
health_check:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
gen_ai_normalizer:
sources:
- name: openinference
remove_originals: true
transform:
trace_statements:
- context: resource
statements:
- set(resource.attributes["deployment.environment.name"], "${env:DEPLOY_ENV}")
redaction:
allow_all_keys: true
blocked_values:
- "sk-[A-Za-z0-9]{20,}" # API keys
- "[0-9]{3}-[0-9]{2}-[0-9]{4}" # US social security numbers
- "4[0-9]{12}(?:[0-9]{3})?" # Visa card numbers
filter:
traces:
span:
- attributes["http.route"] == "/health"
exporters:
otlp_http:
endpoint: https://api.honeycomb.io
headers:
x-honeycomb-team: ${env:HONEYCOMB_API_KEY}
sending_queue:
enabled: true
sizer: items
queue_size: 100000
batch:
flush_timeout: 200ms
min_size: 8192
service:
extensions: [health_check]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, gen_ai_normalizer, transform, redaction, filter]
exporters: [otlp_http]
The processor order matters: memory_limiter goes first so the Collector protects itself before doing any work. gen_ai_normalizer runs before anything that matches on gen_ai.* attributes, so every later stage sees canonical names. Redaction runs after normalization for the same reason; its rules target one set of names instead of one per framework. Note there is no batch processor as batching and retry are handled by the exporter's sending_queue, which is the current recommendation over the standalone processor. Secrets stay in environment variables, never in the config file.
Receive LLM traces over OTLP
The OTLP receiver accepts traces over gRPC on port 4317 and HTTP on port 4318. Most OpenTelemetry SDKs default to one of these, so pointing an application at the Collector is usually a single environment variable, for example OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4318.
Binding to 0.0.0.0 makes the receiver reachable from other containers and pods, which is what you want for a gateway-style deployment. If the Collector runs as a sidecar or on the same host as the application, bind to localhost instead to avoid exposing the port. In Kubernetes, expose the receiver through a Service and point every workload's OTLP endpoint at it.
Enrich and normalize LLM context
If your services use the official GenAI instrumentation libraries, their spans arrive already using the conventions, and the job here is enrichment. The transform processor above stamps deployment.environment.name on every resource. The same pattern adds a team owner, a cost center, or a tenant identifier you're permitted to keep, so you can slice token usage by whatever dimension your organization budgets on.
For services on older, non-OpenTelemetry schemas, the gen_ai_normalizer processor rewrites their attributes onto the conventions. It has built-in mapping tables for OpenInference and OpenLLMetry and accepts user-defined tables for homegrown schemas, with type coercion so string token counts become integers. A sample of what the OpenInference source does:
OpenInference attribute GenAI semantic convention
llm.model_name gen_ai.request.model
llm.provider gen_ai.provider.name
llm.token_count.prompt gen_ai.usage.input_tokens
llm.token_count.completion gen_ai.usage.output_tokens
session.id gen_ai.conversation.id
With remove_originals: true the old names are deleted after mapping, so dashboards only ever see one name. For shapes a mapping table can't express, such as deriving gen_ai.operation.name from span name patterns, pair it with the transform processor.
The normalizer also stamps the scope schema URL it targets, which pays off later. Because the GenAI conventions still evolve, attribute names will move under you. For example, gen_ai.usage.prompt_tokens becoming gen_ai.usage.input_tokens is a real past example. The schema processor uses those schema URLs to translate telemetry between convention versions, and I've written about managing semantic convention migrations with the Collector in depth.
Protect sensitive LLM data before export
This is the section I'd get right before anything ships. Prompts and completions are user content. Tool arguments can carry credentials. Retrieval chunks quote your internal documents. The Collector is the last point where you control that data, so I think the redaction policy belongs here, not in application code you'd have to audit service by service.
The redaction processor supports three complementary controls. blocked_values masks matching substrings wherever they appear, so a card number or air miles reference inside gen_ai.input.messages is masked while the rest of the conversation stays readable, and the same patterns catch API keys that leak into any other attribute. blocked_key_patterns masks the values of matching attribute keys entirely, the strictest option shown commented out in the config above. And hash_function (including hmac-sha256) replaces matches with a stable hash instead of a fixed mask, useful for user identifiers where you need to correlate activity without exporting the raw value.
How far to go is your choice, and the tradeoff is what Honeycomb can show you afterwards. Keeping message content with targeted value masking, as the config above does, preserves conversation debugging and Agent Timeline's message view while known-sensitive patterns never leave your network. Masking gen_ai.input.messages and gen_ai.output.messages entirely protects the most, but conversation-centric workflows lose the messages themselves, so you can see that an agent misbehaved without seeing what it said. If policy says an attribute must never leave the environment, delete it with the transform processor's delete_key rather than masking it. Whatever you choose, it happens inside your network boundary, before export. Honeycomb's guide to removing sensitive information before export covers more best practices.
Control LLM telemetry volume and reliability
memory_limiter protects the Collector from payload spikes, and the exporter's sending_queue batches and retries so transient network failures don't drop data. The filter processor removes traffic with no analytical value, health checks in the example, and the same OTTL syntax can drop zero-token spans or verbose sub-steps you never query.
For sampling, the usual goal is to keep the signal and sample the noise. For GenAI data, our recommendation today is not to sample it at all. Honeycomb's sampling guidance for GenAI data explains why: a conversation loses its value when turns are missing, so completeness matters more than volume for this traffic. Filtering noise with the filter processor and trimming verbose attributes give you most of the volume control without breaking conversations, and you can filter unnecessary telemetry beyond the health-check example above. If the surrounding non-LLM services need sampling, the adaptive_tail_sampling processor we contributed to the OpenTelemetry Collector handles those pipelines while your GenAI traffic flows through complete.
Export LLM telemetry to Honeycomb
The otlp_http exporter sends to https://api.honeycomb.io (use https://api.eu1.honeycomb.io for EU teams) with your API key in the x-honeycomb-team header, read from an environment variable. TLS is on by default for this endpoint, so no extra transport config is needed. The sending_queue settings shown earlier handle batching, retry, and backpressure at the export boundary. Check the Honeycomb OTLP documentation for current endpoint and header values, since these are the settings most likely to change over time.
Validate your LLM traces in Honeycomb #
Validation here means confirming the configuration did what you intended, not just that data arrived.
First, confirm normalization landed. Run a query grouped by gen_ai.request.model. If every service appears under canonical model names regardless of which instrumentation it uses, the normalizer and your native instrumentation agree. A COUNT grouped by gen_ai.operation.name should similarly show values such as chat and execute_tool rather than framework-specific span kinds.
Second, confirm redaction applied. Search for a pattern your policy blocks, an API key prefix or a test card number you deliberately sent through. Zero results means the policy held. Spot-check a span and verify its Gen AI tab renders the message with the masked pattern in place of the sensitive value, and the rest of the message intact.
Third, confirm the numbers are usable. SUM(gen_ai.usage.input_tokens) and SUM(gen_ai.usage.output_tokens) grouped by model give you the token spend view, and opening any agent trace should show tool and model spans correctly parented under the workflow span, with deployment.environment.name present from the enrichment stage.
Finally, open a conversation in Honeycomb's Agent Timeline. It's built directly on the GenAI conventions, so it doubles as a conformance check: if your conversation turns, agents, and tool calls render on the timeline, your pipeline is emitting the right shape. The attributes it relies on most are gen_ai.conversation.id on every span of a run and gen_ai.agent.name on each agent's spans, and the instrumenting AI agents guide covers what to send. Both are worth confirming survive your redaction and filter rules.
Troubleshooting common OpenTelemetry Collector problems #
- No traces arriving: Check the app's OTLP endpoint and port (4317 gRPC / 4318 HTTP), receiver binding, and Collector logs for receiver startup errors.
- Authentication errors: The
HONEYCOMB_API_KEYenv var isn't set in the Collector's environment, or the key belongs to the wrong environment. - Missing
gen_ai.*attributes: The normalizer's source doesn't match the instrumentation. Inspect a raw span's attribute names and add a user-defined mapping table if needed. - Redaction not applying: Check processor order. Redaction must run after the normalizer, and rules written against pre-normalization attribute names silently match nothing.
- Data being dropped:
memory_limiterunder pressure or thesending_queueoverflowing. Check the Collector's own metrics and raise limits or queue size. - Unknown component on startup: Your distribution doesn't include that processor. Use the contrib or Honeycomb distribution, or build a custom one with the components you need.
OTel Collector best practices for production LLM systems #
- Keep secrets in environment variables and the config in version control, and test processor rules against sample payloads before deploying. A redaction rule that almost matches won't be noticed until someone audits the exported data.
- Monitor the Collector itself. Send its internal metrics to Honeycomb and alert on queue growth, dropped spans, and memory pressure.
- Re-validate redaction on every instrumentation change. New SDK versions add new attributes, and a new content-bearing attribute is a leak until your policy covers it.
- Size memory realistically for message payloads, which are far larger than typical span attributes, and review your pipeline when the GenAI semantic conventions publish a new version.
For a worked production example of these practices in a different domain, see this OpenTelemetry Collector implementation example.
Start observing LLM workflows with Honeycomb #
Native GenAI instrumentation (or the normalizer for older schemas) gets everything onto one set of attribute names, redaction keeps user content inside your boundary, filtering keeps volume proportional to value, and the OTLP exporter delivers it. From there, Honeycomb's querying does the interesting work, from token spend by model to debugging a single agent conversation. Start sending OpenTelemetry data to Honeycomb and see what your LLM workflows are actually doing.
Want to try it with your own traffic? Sign up for a free Honeycomb account and start querying, or book office hours with one of our Developer Advocates if you want a hand getting set up.