{"slug": "ai-observability-monitor-llm-cost-tokens-latency", "title": "AI Observability - Monitor LLM Cost, Tokens, Latency", "summary": "SigNoz has added an AI Observability section that reads OpenTelemetry GenAI attributes on spans to report LLM cost, token usage, latency, errors, time to first token (TTFT), and tool calls. The Overview dashboard surfaces stat cards for total cost in USD, total tokens, LLM calls, p95 LLM latency, LLM error rate, and p95 TTFT, plus panels for cost and tokens, per-trace usage, and LLM latency and errors. Cost panels require a pricing rule per model; SigNoz Cloud ships default prices for common models, while open-source SigNoz users must add prices themselves, and self-hosted deployments with a custom collector config must add the signozspanmapper and signozllmpricing processors to the traces pipeline.", "body_md": "AI Observability is a section in SigNoz for LLM and agent workloads. It reads the OpenTelemetry GenAI attributes on your spans and shows cost, token usage, latency, errors, time to first token (TTFT), and tool calls.\n\nUse this page to open AI Observability, read the **Overview** dashboard, and make sure that your spans carry the attributes that the dashboard needs.\n\n## Prerequisites\n\n- An LLM or agent application that sends traces to SigNoz with OpenTelemetry GenAI attributes. For setup steps for OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks, see [LLM Observability](https://signoz.io/docs/llm-observability/) .\n- For cost panels, a pricing rule for each model. SigNoz Cloud includes default prices for common models. In open-source SigNoz, add the price for each of your models. To add or change a price, see [Model pricing](https://signoz.io/docs/ai-observability-model-pricing/) .\n- Self-hosted SigNoz with a custom collector configuration: add the `signozspanmapper` and`signozllmpricing` processors to the traces pipeline. See[Add the AI observability processors](https://signoz.io/docs/operate/migration/upgrade-0-143/#add-the-ai-observability-processors-to-a-custom-collector-config) .\n\n## Open AI Observability\n\n1. In the side navigation, open the **More** menu.\n2. Select **AI Observability** .\n\nSigNoz opens the **Overview** tab. The tab bar at the top has four tabs:\n\n- **Overview** : the dashboard that this page describes.\n- **Explorer** : query LLM and AI trace data. See[AI Observability Explorer](https://signoz.io/docs/ai-observability-explorer/) .\n- **Model pricing** : set the price for each model. See[Model pricing](https://signoz.io/docs/ai-observability-model-pricing/) .\n- **Attribute Mapping** : map attributes from other instrumentation libraries to the GenAI names. See[Attribute Mapping](https://signoz.io/docs/ai-observability-attribute-mapping/) .\n\n## The Overview dashboard\n\nThe **Overview** tab shows the **AI Observability Overview** dashboard. SigNoz creates and updates this dashboard for you, so you cannot edit it. To start an alert from a panel, select **Create Alerts** in the menu of that panel.\n\nA row of stat cards is at the top:\n\n| Card | What it shows | \n|---|---|\n| **Total cost** | The total LLM cost, in USD. | \n| **Total tokens** | Input tokens plus output tokens. | \n| **LLM calls** | The number of spans that have `gen_ai.request.model` . | \n| **LLM latency (p95)** | The p95 duration of LLM calls. | \n| **LLM error rate** | The percentage of LLM calls with an error status. | \n| **TTFT (p95)** | The p95 time to first token, in seconds. | \n\nBelow the cards, the panels are in four sections.\n\n### Cost & tokens\n\n| Panel | What it shows | \n|---|---|\n| **Cost over time** | LLM cost by model. | \n| **Token usage over time** | Input, output, cache-read, and cache-write tokens. | \n| **Cost breakdown** | Cost split into input, output, cache-read, and cache-write token cost. | \n| **Cost per LLM call** | The average cost of one LLM call, by model. | \n| **Prompt cache hit ratio** | Cache-read tokens as a share of all input tokens. | \n| **Cost by service** | LLM cost and call count for each service. | \n| **Cost by provider** | LLM cost and call count for each provider. | \n\n### Per-trace usage\n\nAn AI trace is a trace that has at least one LLM, tool, or agent span.\n\n| Panel | What it shows | \n|---|---|\n| **AI traces over time** | The number of AI traces in each interval. | \n| **Calls per AI trace** | The average number of LLM calls and tool calls in a trace. | \n| **Cost per AI trace** | The average and p95 of the total LLM cost of a trace. | \n| **Tokens per AI trace** | The average and p95 of the total tokens of a trace. | \n\n### LLM latency & errors\n\n| Panel | What it shows | \n|---|---|\n| **LLM latency by model** | The p50, p90, p95, and p99 latency and the call count for each model. | \n| **LLM latency over time** | The p95 latency by model. | \n| **Slowest LLM call per trace** | The p50 and p95 of the longest LLM call in a trace. | \n| **Time to first token** | The p95 TTFT by model. | \n| **LLM errors over time** | LLM calls with an error status, by model. | \n| **Error rate by model** | Errored calls, total calls, and the error percentage for each model. | \n| **Finish reasons** | LLM calls by finish reason. If the share of `length` or`max_tokens` goes up, more responses are cut off. | \n\n### Tool calls\n\n| Panel | What it shows | \n|---|---|\n| **Tool call rate** | Tool calls per second, by tool. | \n| **Tool error rate** | The percentage of tool calls with an error status. | \n| **Tool duration** | The p50 and p95 duration of tool calls. | \n| **Top tools** | The most called tools, with the error count and the median duration. | \n| **Top span names** | The most frequent GenAI span names. | \n\n## Filter the dashboard\n\nThe variable bar at the top of the dashboard has four filters. Each filter lets you select more than one value, or **All**.\n\n| Filter | Attribute | \n|---|---|\n| **model** | `gen_ai.request.model` | \n| **provider** | `gen_ai.provider.name` | \n| **environment** | `deployment.environment` (resource attribute) | \n| **service_name** | `service.name` (resource attribute) | \n\nThe **model** and **provider** filters apply to the LLM and per-trace panels only. Tool spans do not carry a model, so the **Tool calls** panels and **Calls per AI trace** use only the **environment** and **service_name** filters.\n\n## Required attributes\n\nThe dashboard reads these attributes from your spans. A panel stays empty until your spans carry the attributes that it uses. Most names come from the [OpenTelemetry GenAI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/).\n\n| Attribute | Used by | \n|---|---|\n| `gen_ai.request.model` | All LLM panels and the **model** filter. SigNoz counts a span as an LLM call when it has this attribute. | \n| `gen_ai.provider.name` | **Cost by provider** and the**provider** filter. | \n| `gen_ai.usage.input_tokens` ,`gen_ai.usage.output_tokens` | Token panels and cost. | \n| `gen_ai.usage.cache_read.input_tokens` ,`gen_ai.usage.cache_creation.input_tokens` | Cache token panels, cache cost, and **Prompt cache hit ratio** . | \n| `gen_ai.server.ttft` | **TTFT (p95)** and**Time to first token** . Streaming instrumentations, such as OpenLIT, send this attribute in seconds. | \n| `gen_ai.response.finish_reasons` | **Finish reasons** . | \n| `gen_ai.tool.name` | All **Tool calls** panels. SigNoz counts a span as a tool call when it has this attribute. | \n| `gen_ai.agent.name` | Marks agent spans for the **Per-trace usage** panels. | \n| `deployment.environment` ,`service.name` | The **environment** and**service_name** filters, and**Cost by service** . | \n\nError panels use the error status of the span, not an attribute.\n\nYou do not send the cost attributes. SigNoz calculates cost from the token counts and the model price when it receives a span, and it writes the result to `signoz.gen_ai.usage.tokens.cost` and to one cost attribute for each token type.\n\nIf your instrumentation uses other attribute names, map them to these names in [Attribute Mapping](https://signoz.io/docs/ai-observability-attribute-mapping/).\n\n## Limitations\n\n- You cannot edit or clone the Overview dashboard. To change a panel, build it in your own dashboard in [Dashboards](https://signoz.io/docs/userguide/manage-dashboards/) .\n- The earlier in-product `/llm-observability/*` routes no longer exist. Change saved links to`/ai-observability/*` .\n\n## Related\n\n- [AI Observability Explorer](https://signoz.io/docs/ai-observability-explorer/) : query LLM and AI trace data.\n- [Model pricing](https://signoz.io/docs/ai-observability-model-pricing/) : set the price for each model.\n- [Attribute Mapping](https://signoz.io/docs/ai-observability-attribute-mapping/) : map non-standard attributes to GenAI names.\n- [LLM Observability integrations](https://signoz.io/docs/llm-observability/) : instrument OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks.\n\n## Get Help\n\nIf you need help with the steps in this topic, please reach out to us on [SigNoz Community Slack](https://signoz.io/slack/). If you are a SigNoz Cloud user, please use in product chat support located at the bottom right corner of your SigNoz instance or contact us at [cloud-support@signoz.io](mailto:cloud-support@signoz.io).", "url": "https://wpnews.pro/news/ai-observability-monitor-llm-cost-tokens-latency", "canonical_source": "https://signoz.io/docs/ai-observability", "published_at": "2026-09-29 00:00:00+00:00", "updated_at": "2026-09-29 12:48:40.006460+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops", "large-language-models", "ai-agents", "developer-tools"], "entities": ["SigNoz", "OpenTelemetry", "OpenAI", "Anthropic", "LiteLLM", "CrewAI", "signozspanmapper", "signozllmpricing"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-observability-monitor-llm-cost-tokens-latency", "markdown": "https://wpnews.pro/news/ai-observability-monitor-llm-cost-tokens-latency.md", "text": "https://wpnews.pro/news/ai-observability-monitor-llm-cost-tokens-latency.txt", "jsonld": "https://wpnews.pro/news/ai-observability-monitor-llm-cost-tokens-latency.jsonld"}}