cd /news/artificial-intelligence/reduce-llm-spend-with-semantic-routi… · home topics artificial-intelligence article
[ARTICLE · art-67116] src=agentgateway.dev ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Reduce LLM Spend with Semantic Routing

A cost-based semantic routing demo combining vLLM Semantic Router (vSR) with agentgateway reduced LLM spend by 35.4% on a 50-prompt Go and Rust coding dataset, according to a developer evaluation published July 16, 2026. The routed lane cost $0.591444 versus $0.914840 for the always-expensive lane, with 90% expected-tier agreement (45 of 50 prompts) and lower p50 latency (5.22 vs 7.55 seconds).

read5 min views1 publishedJul 20, 2026

I spend a lot of time in Go and Rust, where a quick “one more thing” in a long coding chat can turn a tidy refactor into an hour of debugging. After the hard part is over, routine tests and follow-up questions can still burn premium-model credits. Sending every request to a cheaper model is not the answer: distributed-systems design, correctness work, and difficult debugging need more capable help.

This example combines vLLM Semantic Router (vSR) with agentgateway to send routine Go and Rust coding tasks to gpt-5.4-nano

and escalate complex correctness, distributed-systems, and deep-debugging work to gpt-5.5

. The goal is a reproducible cost reduction without quietly routing everything to the cheapest tier.

The client sends an OpenAI-compatible request with model: "auto"

. An agentgateway policy sends it to vSR, which evaluates semantic, complexity, keyword, context, and structure signals before selecting a model. Agentgateway forwards the request and records the outcome. Cost is not a classifier input; it is measured after the request, keeping the policy explainable.

auto

opts into that automatic selection. A client can instead request gpt-5.4-nano

or gpt-5.5

directly to force a tier and bypass automatic routing; the request still passes through agentgateway. Organizations that require automatic routing can use a request-body transformation to rewrite every request to auto

.

My company needs prices, not just token counts, to make a budget decision. Agentgateway loads a model cost catalog and exposes realized request cost in metrics, logs, and traces. The model cost catalog guide explains how to generate and load the catalog with agctl

.

For this evaluation, I use the integrated OpenTelemetry stack. Prometheus, Loki, Tempo, the OpenTelemetry Collector, and Grafana let engineering and finance inspect the same model-selection event.

I built the cost-based semantic routing demo around a compact 50-prompt Go and Rust dataset. Half of the prompts are routine implementation or test work; the other half involve correctness or distributed-systems problems. I send every prompt through two lanes:

Lane Purpose
routed
vSR selects gpt-5.4-nano or gpt-5.5 .
always_expensive
Every request uses gpt-5.5 ; the cost baseline.

Each run has an ID on every request, so its cost query excludes unrelated cluster traffic. I deliberately leave out a forced-cheap lane: the useful comparison is whether routing saves money while still escalating complex work.

Each prompt has an expected tier in the checked-in dataset: 25 routine prompts expect gpt-5.4-nano

, and 25 complex prompts expect gpt-5.5

. Expected-tier agreement is the share of requests where vSR selected that labelled tier. It checks the routing policy, not the quality of the model’s answer.

In a clean run on July 16, 2026, all 100 primary requests completed with exact model-catalog lookups. vSR selected gpt-5.4-nano

for 30 requests and gpt-5.5

for 20. Catalog-priced agentgateway metrics reported $0.591444

for the routed lane versus $0.914840

for the always-expensive lane, a 35.4% cost reduction. Routed p50 latency was 5.22 seconds, compared with 7.55 seconds for the always-expensive lane. P95 was 19.09 seconds for routed traffic and 19.61 seconds for the always-expensive lane.

The chart reached 90% expected-tier agreement (45 of 50 prompts). All 25 routine prompts stayed on nano, and 20 of 25 complex prompts escalated to GPT-5.5. This is not a claim of perfect model selection, but it shows the savings did not come from a blanket downgrade: 40% of routed requests still used the expensive model. It is a policy sanity check, not an answer-quality benchmark or a substitute for application task-success and user-feedback data.

The default dataset uses independent requests. For long-running coding agents, the optional cache-transition evaluation uses two ordered ten-turn Go and Rust conversations with a stable prefix. It reports provider-observed cache reads by model transition without mixing cache behavior into the primary evaluation.

The demo automates the setup work I would otherwise repeat by hand. It uses the agentgateway semantic-routing example, creates or reuses a kind cluster, and installs agentgateway, vSR, a model cost catalog, and the OpenTelemetry stack. It checks readiness, routing, catalog-priced metrics, logs, and traces before the primary requests run.

Before running it, review the demo requirements: 12 GiB of Docker memory, 30 GiB of free disk, the required CLIs, an OpenAI API key, and access to both models.

git clone https://github.com/danehans/agentgateway-demos.git
cd agentgateway-demos/cost-based-semantic-routing

export OPENAI_API_KEY='sk-...'
./demo.sh all --yes

The default run sends 106 billable requests: two routing probes, four smoke-test requests, and 100 primary requests. It writes request-level JSONL, an evaluation-scoped Prometheus report, and an SVG chart under results/

.

When I want to change the policy, I edit the fetched vSR values, redeploy with ./demo.sh router

, and run ./demo.sh eval --yes

again.

Codex lets me choose a model manually, but a user-level profile can instead send every request to a gateway such as agentgateway with the stable auto

model name:

model = "auto"
model_provider = "agentgateway"

[model_providers.agentgateway]
name = "Corporate agentgateway"
base_url = "https://my.corp.agentgateway.com/v1"
wire_api = "responses"

I start Codex with codex --profile agentgateway

. It sends OpenAI Responses API traffic to agentgateway; vSR selects the tier and agentgateway records the decision and realized cost.

Current Codex emits an unknown model auto

fallback-metadata warning. Routing still works, but Codex warns that the fallback metadata may degrade behavior.

The profile keeps credentials out of the configuration file. Codex documents custom model providers and user-level configuration profiles.

Semantic routing is not a magic cost switch. Start with an explainable policy, compare it with always using the expensive model, confirm it uses both tiers, and tune it against task completion, feedback, retries, and escalation rates. Agentgateway supplies the accounting and observability; vSR makes the semantic decision.

This example is intentionally static: it uses configured keywords, semantic candidates, weights, and thresholds rather than learning from application outcomes. Those signals need periodic review as traffic changes. The agentgateway community and vSR project are continuing work to make semantic routing easier to tune, operate, and measure against real application traffic, so stay tuned.

After validating a tiering policy, use agentgateway cost controls to attribute usage with virtual keys, observe realized spend by model and consumer, and enforce budget or spend limits.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @vllm semantic router 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reduce-llm-spend-wit…] indexed:0 read:5min 2026-07-20 ·