AI Observability - Monitor LLM Cost, Tokens, Latency SigNoz has added an AI Observability section that reads OpenTelemetry GenAI attributes on spans to report LLM cost, token usage, latency, errors, time to first token (TTFT), and tool calls. The Overview dashboard surfaces stat cards for total cost in USD, total tokens, LLM calls, p95 LLM latency, LLM error rate, and p95 TTFT, plus panels for cost and tokens, per-trace usage, and LLM latency and errors. Cost panels require a pricing rule per model; SigNoz Cloud ships default prices for common models, while open-source SigNoz users must add prices themselves, and self-hosted deployments with a custom collector config must add the signozspanmapper and signozllmpricing processors to the traces pipeline. AI Observability is a section in SigNoz for LLM and agent workloads. It reads the OpenTelemetry GenAI attributes on your spans and shows cost, token usage, latency, errors, time to first token TTFT , and tool calls. Use this page to open AI Observability, read the Overview dashboard, and make sure that your spans carry the attributes that the dashboard needs. Prerequisites - An LLM or agent application that sends traces to SigNoz with OpenTelemetry GenAI attributes. For setup steps for OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks, see LLM Observability https://signoz.io/docs/llm-observability/ . - For cost panels, a pricing rule for each model. SigNoz Cloud includes default prices for common models. In open-source SigNoz, add the price for each of your models. To add or change a price, see Model pricing https://signoz.io/docs/ai-observability-model-pricing/ . - Self-hosted SigNoz with a custom collector configuration: add the signozspanmapper and signozllmpricing processors to the traces pipeline. See Add the AI observability processors https://signoz.io/docs/operate/migration/upgrade-0-143/ add-the-ai-observability-processors-to-a-custom-collector-config . Open AI Observability 1. In the side navigation, open the More menu. 2. Select AI Observability . SigNoz opens the Overview tab. The tab bar at the top has four tabs: - Overview : the dashboard that this page describes. - Explorer : query LLM and AI trace data. See AI Observability Explorer https://signoz.io/docs/ai-observability-explorer/ . - Model pricing : set the price for each model. See Model pricing https://signoz.io/docs/ai-observability-model-pricing/ . - Attribute Mapping : map attributes from other instrumentation libraries to the GenAI names. See Attribute Mapping https://signoz.io/docs/ai-observability-attribute-mapping/ . The Overview dashboard The Overview tab shows the AI Observability Overview dashboard. SigNoz creates and updates this dashboard for you, so you cannot edit it. To start an alert from a panel, select Create Alerts in the menu of that panel. A row of stat cards is at the top: | Card | What it shows | |---|---| | Total cost | The total LLM cost, in USD. | | Total tokens | Input tokens plus output tokens. | | LLM calls | The number of spans that have gen ai.request.model . | | LLM latency p95 | The p95 duration of LLM calls. | | LLM error rate | The percentage of LLM calls with an error status. | | TTFT p95 | The p95 time to first token, in seconds. | Below the cards, the panels are in four sections. Cost & tokens | Panel | What it shows | |---|---| | Cost over time | LLM cost by model. | | Token usage over time | Input, output, cache-read, and cache-write tokens. | | Cost breakdown | Cost split into input, output, cache-read, and cache-write token cost. | | Cost per LLM call | The average cost of one LLM call, by model. | | Prompt cache hit ratio | Cache-read tokens as a share of all input tokens. | | Cost by service | LLM cost and call count for each service. | | Cost by provider | LLM cost and call count for each provider. | Per-trace usage An AI trace is a trace that has at least one LLM, tool, or agent span. | Panel | What it shows | |---|---| | AI traces over time | The number of AI traces in each interval. | | Calls per AI trace | The average number of LLM calls and tool calls in a trace. | | Cost per AI trace | The average and p95 of the total LLM cost of a trace. | | Tokens per AI trace | The average and p95 of the total tokens of a trace. | LLM latency & errors | Panel | What it shows | |---|---| | LLM latency by model | The p50, p90, p95, and p99 latency and the call count for each model. | | LLM latency over time | The p95 latency by model. | | Slowest LLM call per trace | The p50 and p95 of the longest LLM call in a trace. | | Time to first token | The p95 TTFT by model. | | LLM errors over time | LLM calls with an error status, by model. | | Error rate by model | Errored calls, total calls, and the error percentage for each model. | | Finish reasons | LLM calls by finish reason. If the share of length or max tokens goes up, more responses are cut off. | Tool calls | Panel | What it shows | |---|---| | Tool call rate | Tool calls per second, by tool. | | Tool error rate | The percentage of tool calls with an error status. | | Tool duration | The p50 and p95 duration of tool calls. | | Top tools | The most called tools, with the error count and the median duration. | | Top span names | The most frequent GenAI span names. | Filter the dashboard The variable bar at the top of the dashboard has four filters. Each filter lets you select more than one value, or All . | Filter | Attribute | |---|---| | model | gen ai.request.model | | provider | gen ai.provider.name | | environment | deployment.environment resource attribute | | service name | service.name resource attribute | The model and provider filters apply to the LLM and per-trace panels only. Tool spans do not carry a model, so the Tool calls panels and Calls per AI trace use only the environment and service name filters. Required attributes The dashboard reads these attributes from your spans. A panel stays empty until your spans carry the attributes that it uses. Most names come from the OpenTelemetry GenAI semantic conventions https://opentelemetry.io/docs/specs/semconv/gen-ai/ . | Attribute | Used by | |---|---| | gen ai.request.model | All LLM panels and the model filter. SigNoz counts a span as an LLM call when it has this attribute. | | gen ai.provider.name | Cost by provider and the provider filter. | | gen ai.usage.input tokens , gen ai.usage.output tokens | Token panels and cost. | | gen ai.usage.cache read.input tokens , gen ai.usage.cache creation.input tokens | Cache token panels, cache cost, and Prompt cache hit ratio . | | gen ai.server.ttft | TTFT p95 and Time to first token . Streaming instrumentations, such as OpenLIT, send this attribute in seconds. | | gen ai.response.finish reasons | Finish reasons . | | gen ai.tool.name | All Tool calls panels. SigNoz counts a span as a tool call when it has this attribute. | | gen ai.agent.name | Marks agent spans for the Per-trace usage panels. | | deployment.environment , service.name | The environment and service name filters, and Cost by service . | Error panels use the error status of the span, not an attribute. You do not send the cost attributes. SigNoz calculates cost from the token counts and the model price when it receives a span, and it writes the result to signoz.gen ai.usage.tokens.cost and to one cost attribute for each token type. If your instrumentation uses other attribute names, map them to these names in Attribute Mapping https://signoz.io/docs/ai-observability-attribute-mapping/ . Limitations - You cannot edit or clone the Overview dashboard. To change a panel, build it in your own dashboard in Dashboards https://signoz.io/docs/userguide/manage-dashboards/ . - The earlier in-product /llm-observability/ routes no longer exist. Change saved links to /ai-observability/ . Related - AI Observability Explorer https://signoz.io/docs/ai-observability-explorer/ : query LLM and AI trace data. - Model pricing https://signoz.io/docs/ai-observability-model-pricing/ : set the price for each model. - Attribute Mapping https://signoz.io/docs/ai-observability-attribute-mapping/ : map non-standard attributes to GenAI names. - LLM Observability integrations https://signoz.io/docs/llm-observability/ : instrument OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks. Get Help If you need help with the steps in this topic, please reach out to us on SigNoz Community Slack https://signoz.io/slack/ . If you are a SigNoz Cloud user, please use in product chat support located at the bottom right corner of your SigNoz instance or contact us at cloud-support@signoz.io mailto:cloud-support@signoz.io .