cd /news/ai-infrastructure/ai-observability-monitor-llm-cost-to… · home › topics › ai-infrastructure › article
[ARTICLE · art-141691] src=signoz.io ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

AI Observability - Monitor LLM Cost, Tokens, Latency

SigNoz has added an AI Observability section that reads OpenTelemetry GenAI attributes on spans to report LLM cost, token usage, latency, errors, time to first token (TTFT), and tool calls. The Overview dashboard surfaces stat cards for total cost in USD, total tokens, LLM calls, p95 LLM latency, LLM error rate, and p95 TTFT, plus panels for cost and tokens, per-trace usage, and LLM latency and errors. Cost panels require a pricing rule per model; SigNoz Cloud ships default prices for common models, while open-source SigNoz users must add prices themselves, and self-hosted deployments with a custom collector config must add the signozspanmapper and signozllmpricing processors to the traces pipeline.

by read6 min views1 publishedSep 29, 2026

AI Observability is a section in SigNoz for LLM and agent workloads. It reads the OpenTelemetry GenAI attributes on your spans and shows cost, token usage, latency, errors, time to first token (TTFT), and tool calls.

Use this page to open AI Observability, read the Overview dashboard, and make sure that your spans carry the attributes that the dashboard needs.

Prerequisites #

  • An LLM or agent application that sends traces to SigNoz with OpenTelemetry GenAI attributes. For setup steps for OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks, see LLM Observability .
  • For cost panels, a pricing rule for each model. SigNoz Cloud includes default prices for common models. In open-source SigNoz, add the price for each of your models. To add or change a price, see Model pricing .
  • Self-hosted SigNoz with a custom collector configuration: add the signozspanmapper andsignozllmpricing processors to the traces pipeline. SeeAdd the AI observability processors .

Open AI Observability #

  1. In the side navigation, open the More menu.
  2. Select AI Observability .

SigNoz opens the Overview tab. The tab bar at the top has four tabs:

  • Overview : the dashboard that this page describes.
- **Explorer** : query LLM and AI trace data. See[AI Observability Explorer](https://signoz.io/docs/ai-observability-explorer/) .
- **Model pricing** : set the price for each model. See[Model pricing](https://signoz.io/docs/ai-observability-model-pricing/) .
- **Attribute Mapping** : map attributes from other instrumentation libraries to the GenAI names. See[Attribute Mapping](https://signoz.io/docs/ai-observability-attribute-mapping/) .

The Overview dashboard #

The Overview tab shows the AI Observability Overview dashboard. SigNoz creates and updates this dashboard for you, so you cannot edit it. To start an alert from a panel, select Create Alerts in the menu of that panel.

A row of stat cards is at the top:

Card What it shows
Total cost The total LLM cost, in USD.
Total tokens Input tokens plus output tokens.
LLM calls The number of spans that have gen_ai.request.model .
LLM latency (p95) The p95 duration of LLM calls.
LLM error rate The percentage of LLM calls with an error status.
TTFT (p95) The p95 time to first token, in seconds.

Below the cards, the panels are in four sections.

Cost & tokens

Panel What it shows
Cost over time LLM cost by model.
Token usage over time Input, output, cache-read, and cache-write tokens.
Cost breakdown Cost split into input, output, cache-read, and cache-write token cost.
Cost per LLM call The average cost of one LLM call, by model.
Prompt cache hit ratio Cache-read tokens as a share of all input tokens.
Cost by service LLM cost and call count for each service.
Cost by provider LLM cost and call count for each provider.

Per-trace usage

An AI trace is a trace that has at least one LLM, tool, or agent span.

Panel What it shows
AI traces over time The number of AI traces in each interval.
Calls per AI trace The average number of LLM calls and tool calls in a trace.
Cost per AI trace The average and p95 of the total LLM cost of a trace.
Tokens per AI trace The average and p95 of the total tokens of a trace.

LLM latency & errors

Panel What it shows
LLM latency by model The p50, p90, p95, and p99 latency and the call count for each model.
LLM latency over time The p95 latency by model.
Slowest LLM call per trace The p50 and p95 of the longest LLM call in a trace.
Time to first token The p95 TTFT by model.
LLM errors over time LLM calls with an error status, by model.
Error rate by model Errored calls, total calls, and the error percentage for each model.
Finish reasons LLM calls by finish reason. If the share of length ormax_tokens goes up, more responses are cut off.

Tool calls

Panel What it shows
Tool call rate Tool calls per second, by tool.
Tool error rate The percentage of tool calls with an error status.
Tool duration The p50 and p95 duration of tool calls.
Top tools The most called tools, with the error count and the median duration.
Top span names The most frequent GenAI span names.

Filter the dashboard #

The variable bar at the top of the dashboard has four filters. Each filter lets you select more than one value, or All.

Filter Attribute
model gen_ai.request.model
provider gen_ai.provider.name
environment deployment.environment (resource attribute)
service_name service.name (resource attribute)

The model and provider filters apply to the LLM and per-trace panels only. Tool spans do not carry a model, so the Tool calls panels and Calls per AI trace use only the environment and service_name filters.

Required attributes #

The dashboard reads these attributes from your spans. A panel stays empty until your spans carry the attributes that it uses. Most names come from the OpenTelemetry GenAI semantic conventions.

Attribute Used by
gen_ai.request.model All LLM panels and the model filter. SigNoz counts a span as an LLM call when it has this attribute.
gen_ai.provider.name Cost by provider and theprovider filter.
gen_ai.usage.input_tokens ,gen_ai.usage.output_tokens Token panels and cost.
gen_ai.usage.cache_read.input_tokens ,gen_ai.usage.cache_creation.input_tokens Cache token panels, cache cost, and Prompt cache hit ratio .
gen_ai.server.ttft TTFT (p95) andTime to first token . Streaming instrumentations, such as OpenLIT, send this attribute in seconds.
gen_ai.response.finish_reasons Finish reasons .
gen_ai.tool.name All Tool calls panels. SigNoz counts a span as a tool call when it has this attribute.
gen_ai.agent.name Marks agent spans for the Per-trace usage panels.
deployment.environment ,service.name The environment andservice_name filters, andCost by service .

Error panels use the error status of the span, not an attribute.

You do not send the cost attributes. SigNoz calculates cost from the token counts and the model price when it receives a span, and it writes the result to signoz.gen_ai.usage.tokens.cost and to one cost attribute for each token type.

If your instrumentation uses other attribute names, map them to these names in Attribute Mapping.

Limitations #

  • You cannot edit or clone the Overview dashboard. To change a panel, build it in your own dashboard in Dashboards .
  • The earlier in-product /llm-observability/* routes no longer exist. Change saved links to/ai-observability/* .
- [AI Observability Explorer](https://signoz.io/docs/ai-observability-explorer/) : query LLM and AI trace data.
- [Model pricing](https://signoz.io/docs/ai-observability-model-pricing/) : set the price for each model.
- [Attribute Mapping](https://signoz.io/docs/ai-observability-attribute-mapping/) : map non-standard attributes to GenAI names.
- [LLM Observability integrations](https://signoz.io/docs/llm-observability/) : instrument OpenAI, Anthropic, LiteLLM, CrewAI, and other frameworks.

Get Help #

If you need help with the steps in this topic, please reach out to us on SigNoz Community Slack. If you are a SigNoz Cloud user, please use in product chat support located at the bottom right corner of your SigNoz instance or contact us at cloud-support@signoz.io.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @signoz 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-observability-mon…] indexed:0 read:6min 2026-09-29 · —