{"slug": "agentgateway-v1-6-0-cost-tracking-with-no-catalog-config", "title": "agentgateway v1.6.0: Cost Tracking With No Catalog Config", "summary": "A developer tested agentgateway v1.6.0, released October 2, and found that its built-in LLM pricing catalog delivers per-request cost tracking with no modelCatalog configuration, logging fields such as cost.total=0.000156 for a Claude Sonnet 5 call. The release also adds per-key CEL rate limiting that buckets quota by any request attribute, demonstrated by keying a 3-request-per-60-second limit on the x-api-key header so separate keys never share a bucket.", "body_md": "*Originally published at [webofmike.com](https://webofmike.com/agentgateway-v16-cost-tracking/?utm_source=devto&utm_medium=syndication&utm_campaign=agentgateway-v16-cost-tracking) on 2026-10-06. The demo repo and every command in it were run before publishing.*\n\n[agentgateway v1.6.0](https://github.com/agentgateway/agentgateway/releases/tag/v1.6.0) went GA on October 2. Two features are worth a hands-on look before anything else: a built-in LLM pricing catalog that gives you cost tracking with zero configuration, and per-key CEL rate limiting that buckets quota by any request attribute you pick. I ran both against a live gateway and a real Claude Sonnet 5 backend. Code and captured output are in [themsquared/agw-16-hands-on](https://github.com/themsquared/agw-16-hands-on).\n\nI [wrote up agentgateway's per-API-key budgets back in v1.5.0](https://webofmike.com/agentgateway-per-key-llm-budgets/): a `budgets` list per key, a limit in tokens or USD, and a 429 before the provider ever sees the request. That post also covered the gap underneath it. Budgets only work once the gateway knows what a token costs, and in 1.5 that meant a `modelCatalog` block naming every model and its per-million-token input and output rate, maintained by hand.\n\nv1.6.0 replaces that with a model catalog agentgateway ships and maintains itself. The demo's entire LLM route config:\n\n```\nllm:\n  port: 4000\n  policies:\n    localRateLimit:\n    - type: requests\n      maxTokens: 3\n      tokensPerFill: 3\n      fillInterval: 60s\n      key: request.headers[\"x-api-key\"]\n  models:\n  - name: claude\n    provider: anthropic\n    params:\n      model: claude-sonnet-5\n      apiKey: $ANTHROPIC_API_KEY\n```\n\nNo `modelCatalog` anywhere. The access log still produces full cost data because `claude-sonnet-5` is a model the built-in catalog already knows:\n\n```\nhttp.status=200 gen_ai.usage.input_tokens=8 gen_ai.usage.output_tokens=14\nagw.ai.usage.cost.total=0.000156\ncost.total=0.000156 cost.rate.input=2 cost.rate.output=10\n```\n\n`cost.rate.input` and `cost.rate.output` are USD per million tokens, pulled straight from the catalog. The math checks out: `(8 * 2 + 14 * 10) / 1,000,000 = 0.000156`. Nothing in the config told agentgateway what this model costs; it already knew.\n\nGetting those fields into the access log in the first place is a `frontendPolicies.accessLog.add` block mapping CEL expressions to log keys:\n\n```\nfrontendPolicies:\n  accessLog:\n    add:\n      model.requested: llm.requestModel\n      model.served: llm.responseModel\n      tokens.input: llm.inputTokens\n      tokens.output: llm.outputTokens\n      cost.total: llm.cost.total\n      cost.rate.input: llm.costRates.input\n      cost.rate.output: llm.costRates.output\n```\n\nThose same CEL fields are what you'd read in a `budgets` policy from the 1.5 post, so this isn't a separate feature bolted on. It's the same cost-accounting plumbing, now backed by a catalog instead of a hand-maintained table.\n\nThe second feature in this run is `localRateLimit`, keyed by a CEL expression rather than a fixed value. The config above keys on `request.headers[\"x-api-key\"]`, with a 3-request bucket that refills every 60 seconds. One rule, evaluated per request, buckets independently per header value:\n\n```\n# key A, requests 1-3: all 200\nhttp.status=200 ... cost.total=0.000156 ...\nhttp.status=200 ...\nhttp.status=200 ...\n\n# key A, request 4, same minute\nhttp.status=429 error=\"rate limit exceeded\" reason=RateLimit\n\n# key B, same minute, different header value\nhttp.status=200 ... cost.total=0.000176\n```\n\nKey B never saw key A's limit. There's no second `localRateLimit` entry for it, no restart to pick up a new key. Any CEL expression over the request works as the bucket key, so this generalizes past API keys to things like a JWT claim or a source IP.\n\nOne run of the demo hit a transient upstream failure partway through:\n\n```\nhttp.status=503 error=\"upstream call failed: SendRequest: connection error: peer closed connection without sending TLS close_notify\" reason=UpstreamFailure\n```\n\n`x-ratelimit-remaining` still dropped on that request. The call never reached Anthropic successfully, a `503` is the opposite of a billable response, but it still counted as one of the 3 admitted requests in the bucket. `localRateLimit` counts requests agentgateway admits, not requests that succeed upstream. If you're sizing `maxTokens` close to real traffic, budget headroom for upstream flakiness, because a bad backend day eats your quota exactly like a good one.\n\n`AgentgatewayModel` CRD, now on by default in the Helm chart, and K8s-native session affinity are cluster-side features this standalone run doesn't exercise.`remoteRateLimit` is a different policy, for quota shared across replicas. `localRateLimit` buckets live in the single proxy instance that created them, which is exactly what makes this demo's single-container setup representative of the behavior.\nRequirements: Docker and an `ANTHROPIC_API_KEY`. Tested on macOS (Apple silicon) against `cr.agentgateway.dev/agentgateway:v1.6.0`.\n\n```\ngit clone https://github.com/themsquared/agw-16-hands-on.git\ncd agw-16-hands-on\nexport ANTHROPIC_API_KEY=sk-ant-...\n./run-demo.sh\n```\n\nTo check the config against the v1.6.0 schema without sending any traffic:\n\n```\ndocker run --rm -v \"$PWD/config/config.yaml:/config/config.yaml:ro\" \\\n  -e ANTHROPIC_API_KEY=dummy-for-validate \\\n  cr.agentgateway.dev/agentgateway:v1.6.0 -f /config/config.yaml --validate-only\n```\n\nTear down with `docker rm -f agw16-demo`.\n\nGoing from v1.5.0 to v1.6.0, the same cost-and-quota problem from the [per-key budgets post](https://webofmike.com/agentgateway-per-key-llm-budgets/) now needs less from you: no catalog to maintain, and a quota rule that keys itself by request content instead of being written once per key. The gotcha is the same shape either version: a limiter counts what it admits, not what succeeds, and that is worth checking against your own traffic patterns before you pick a `maxTokens` value. Repo and full captured output: [themsquared/agw-16-hands-on](https://github.com/themsquared/agw-16-hands-on).\n\n**Does agentgateway v1.6.0 need a model catalog configured for cost tracking?**\n\nNo. v1.6.0 ships a built-in model catalog, so a route to a known model like claude-sonnet-5 produces llm.cost.total and the input/output cost rates automatically, with no modelCatalog block anywhere in the config. Earlier versions required declaring rates by hand for every model in use.\n\n**How does agentgateway's per-key CEL rate limiting work?**\n\nA single localRateLimit policy keyed on an expression like request.headers[\\\n\n**Does a failed upstream request still count against an agentgateway rate limit bucket?**\n\nYes. A request that fails upstream, such as a dropped TLS connection, still consumes a token from the bucket. localRateLimit counts requests admitted at the gateway, not successful upstream responses, so a flaky backend can eat into a tight per-key quota without a single call completing.\n\n**Is agentgateway's remoteRateLimit the same as localRateLimit?**\n\nNo. localRateLimit buckets live in memory on the single gateway instance that created them, which is exactly the single-container setup this demo runs. remoteRateLimit is a separate policy that shares one quota across replicas, which matters once more than one gateway instance sits behind a load balancer and you need one limit enforced across all of them.\n\n*Canonical version, with machine-readable markdown at `https://webofmike.com/agentgateway-v16-cost-tracking/index.md`: [https://webofmike.com/agentgateway-v16-cost-tracking/](https://webofmike.com/agentgateway-v16-cost-tracking/)*", "url": "https://wpnews.pro/news/agentgateway-v1-6-0-cost-tracking-with-no-catalog-config", "canonical_source": "https://dev.to/webofmike/agentgateway-v160-cost-tracking-with-no-catalog-config-374c", "published_at": "2026-10-06 18:11:12+00:00", "updated_at": "2026-10-06 18:18:24.848747+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "developer-tools", "mlops"], "entities": ["agentgateway", "agentgateway v1.6.0", "Claude Sonnet 5", "Anthropic", "themsquared/agw-16-hands-on", "webofmike.com"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agentgateway-v1-6-0-cost-tracking-with-no-catalog-config", "markdown": "https://wpnews.pro/news/agentgateway-v1-6-0-cost-tracking-with-no-catalog-config.md", "text": "https://wpnews.pro/news/agentgateway-v1-6-0-cost-tracking-with-no-catalog-config.txt", "jsonld": "https://wpnews.pro/news/agentgateway-v1-6-0-cost-tracking-with-no-catalog-config.jsonld"}}