Part 5: Operating an LLM system: observability, cost, routing, and the platform underneath A six-level maturity model for running LLM systems in production defines Level 5 as the operability layer, specifying a decision-layer observability schema built on RED metrics plus AI-specific signals such as agent_decisions_total, agent_auto_execution_ratio, guardrail_blocks_total, judge_disagreements_total, llm_tokens_total, llm_cost_usd_total, and shadow_eval_pass_ratio, with every series labeled by tenant_id and capability. The model routes all model traffic through a single chokepoint tagged with one decision_id, so the domain never knows which vendor answered, and warns that a per-tenant_id label on histograms becomes a cardinality bomb past a few hundred tenants. The author argues the asymmetry that an LLM system going wrong can take wrong actions at scale, fast, makes operability the difference between a system you can run and one you have to hope about. This is Level 5 of a six-level maturity model for running LLM systems in production. The earlier levels got the system working and correct . Level 5 is about operability : can you actually run this thing day to day, see what it’s deciding, control what it costs, survive a provider outage, stop it in seconds when it misbehaves — and is the platform underneath shaped to support all of that? A traditional service that goes wrong returns a 500. An LLM system that goes wrong can take wrong actions at scale, fast . That asymmetry is why operability isn’t a nice-to-have here — it’s the difference between a system you can run and one you have to hope about. This post is the concrete spec for the operability layer: the observability signals, the cost controls, the routing and failover, the kill switch, the identity model, and the infrastructure spine that holds it together. It’s the longest post in the series, because it’s the layer with the most moving parts. The unifying idea, though, is small : one chokepoint your model traffic flows through, one decision id that threads everything, and a domain that never knows which vendor answered. Standard observability — request rate, latency, error rate, CPU — tells you the service is up. It tells you nothing about whether the agent is doing its job well. You need a second layer of signals specific to AI, plus the schemas and thresholds to make them actionable. The model: RED + a decision layer Keep your usual RED metrics Rate, Errors, Duration for the service. Add a decision layer that treats every agent decision as a first-class, measurable event. Names below follow Prometheus conventions total counters, bucket histograms , a nd every series carries tenant id and capability as labels omitted in the table for brevity — assume them everywhere . One caveat: a per- tenant id label on histograms is a cardinality bomb past a few hundred tenants — keep tenant counts bounded, or move per-tenant rollups to recording rules / exemplars beyond that. metric type key labels what it answers --------------------------------- --------- ---------------------------------------------------------------------- ------------------------------------------------- agent decisions total counter outcome auto / hitl recommended / hitl required / reject / abstain volume + the auto/human split agent auto execution ratio gauge — % handled without a human agent decision confidence histogram — distribution of composed confidence agent decision duration seconds histogram node latency per graph node guardrail blocks total counter layer input / output / pii , rule what's being blocked, where judge invocations total counter — how often the judge ran the ratio's denominator judge disagreements total counter — judge overruled the primary llm tokens total counter direction in / out , model token consumption llm cost usd total counter model spend, rolled up by tenant/capability llm call duration seconds histogram model , outcome model latency, separate from service shadow eval pass ratio gauge slice live quality on sampled prod traffic human override rate gauge — Four families fall out of that: decision volume, split, confidence, latency , safety guardrail blocks, judge disagreement , cost/perf tokens, USD, model latency , quality shadow eval, overrides . If you only instrument four things, make them auto execution ratio, guardrail blocks total , llm cost usd total, and human override rate — they cover behavior, safety, money, and quality. Structured logging: three tiers Don’t dump everything into one log stream — tier by sensitivity, because some of this is regulated data and most alerts only need tier 1. - Tier 1 — operational safe anywhere, high volume : timestamp, level, decision id, tenant id , capability, node , duration ms , outcome. This is what your alerts query. - Tier 2 — decision metadata inside your trust boundary; the audit-adjacent record : model, prompt version , composed confidence, routing decision, judge result, guardrail actions, an inputs hash , and a redacted PII-free summary. - Tier 3 — regulated/raw encrypted, access-controlled, never in your general log store : the sensitive payload, if you retain it at all — usually you store a hash + redacted summary in tier 2 and skip tier 3 entirely. The rule: tier 1 and 2 are queryable by engineers; tier 3 is a vault. A decision id threads all three — and the trace below — so you can reconstruct any single decision end-to-end. Trace the decision graph A normal trace shows the HTTP request across services. An agent trace should also span the internal graph so you can see which step failed and how long each took: span: agent.decision attrs: decision id, tenant id, capability, outcome, confidence ├─ span: entry attrs: identity, auth type ├─ span: context load attrs: slices loaded, cache hit ├─ span: llm decision attrs: model, prompt version, tokens in, tokens out, cost usd ├─ span: output guardrail attrs: blocks , pii actions ├─ span: judge attrs: ran, model, agreed only when sampled in ├─ span: routing attrs: decision=auto|hitl|reject, threshold └─ span: exit attrs: ledger entry id Now “why did decision X take 4 seconds / get rejected?” is a single trace lookup. Alert on symptoms, with real thresholds Alert on what hurts the user or business, not on causes. A starter set: alert condition example severity --------------------- ------------------------------------------------------------------------------ ------------ AutoExecutionSwing auto execution ratio moves 15% vs 7-day baseline warning GuardrailBlockSpike rate guardrail blocks total 5m 3× trailing hr critical JudgeDisagreementHigh rate judge disagreements total 1h / rate judge invocations total 1h 0.15 critical CostCeiling increase llm cost usd total 1h budget/24 warning→page ShadowEvalDrop shadow eval pass ratio < 0.95 critical ModelLatencyP99 llm call duration seconds p99 8s for 10m Page on the safety and quality ones; the rest open a ticket. And set SLOs that mean something for a decision system: decision availability ≥ 99.9% of decisions return a proposal — failing to decide is the real outage, not a 5xx , quality shadow-eval ≥ baseline − 2%, rolling 7d , latency p95 excluding human time under your interaction budget , and cost USD per 1k decisions within ±20% of plan . Three observability anti-patterns worth naming: one log stream for everything tier-3 leaks into searchable logs — a compliance problem , self-reported model confidence as a metric it’s miscalibrated; track composed confidence and validate against outcomes , and no decision id every investigation becomes archaeology . LLM cost has a nasty property: it’s invisible until the invoice, and it scales with things you’re not watching — tokens per call, calls per request, retries, a chatty prompt someone added. Teams discover their unit economics are upside down only after they’ve shipped. The fix isn’t a cheaper model; it’s a handful of structural controls that make cost observable and bounded — and they all hang off the same chokepoint. Measure where you enforce You can’t control what you can’t see, and you can’t see spend when model calls happen all over the codebase. Route every call through one gateway. That component is where you measure and enforce : php def complete self, req: ChatRequest - ChatResponse: rate limiter.check req.tenant id, est tokens req enforce estimate now; settle actuals below resp = self. adapter.complete req metrics.incr "llm tokens total", resp.tokens in, direction="in", model=req.model, tenant id=req.tenant id metrics.incr "llm tokens total", resp.tokens out, direction="out", model=req.model, tenant id=req.tenant id metrics.incr "llm cost usd total", cost req.model, resp , model=req.model, tenant id=req.tenant id, capability=req.capability return resp Now cost is attributable per tenant × capability × model — which is how you find where the money goes usually one chatty capability or one oversized prompt instead of vaguely “using less AI.” The levers, in order of payoff lever mechanism typical impact ------------------------ ------------------------------------------------------------ ---------------------------------------------------------------------- don't call the model route easy/deterministic cases through code often the biggest — a lot of "AI cost" is the model doing a rule's job batch one call for many items vs N calls fewer round-trips, cheaper per item cache memoize deterministic results embeddings, repeated lookups the cheapest call is the one you skip right-size the model cheap model for simple decisions, capable for hard ones big — most traffic is simple, most cost is the premium model trim the prompt load only needed context; kill "just in case" preamble recurring tax paid on every call Right-sizing is worth making explicit, because it feeds directly into routing next section : if a cheap model passes evals for a slice, route it there — 5–20× cheaper per call is common. Make the math visible cost = tokens in/1k × price in + tokens out/1k × price out rather than guessing. Budgets, rate limits, and the silent multipliers Cap spend and rate per tenant, with state shared across replicas — a stateless worker pool can’t enforce a per-tenant limit from local memory, since each replica would allow its own full fraction: RATE = { "default": { "tokens per min": 200 000, "usd per day": 50 } } derive from real prices def check tenant id, est tokens : limits = RATE FOR tenant id add the ESTIMATED token count, not 1 — counting calls against a token budget never trips if redis.incr window f"tok:{tenant id}", by=est tokens, ttl=60 limits "tokens per min" : raise RateLimited tenant id spent = redis.get float f"usd:{tenant id}:{today }" written by the gateway's post-call cost path if spent limits "usd per day" : raise BudgetExceeded tenant id 100%: hard stop if spent 0.8 limits "usd per day" : alert tenant id, "80% of daily budget" 80%: warn Also set a platform-wide ceiling — N tenants × per-tenant cap has no aggregate bound otherwise. And watch the two things that quietly 10× a bill: retry storms a flaky validation re-calling the model and unbounded tool loops an agent that keeps going . Cap both, and emit llm retries total{reason} so a storm is visible. LLM cost isn’t fundamentally high; it’s fundamentally unmonitored . Monitor it at the chokepoint and the bill stays proportional to value. Cost and latency share a root cause: a pipeline that takes minutes usually isn’t slow because of the model. It’s an O n² loop, a per-record round-trip, or serial stages — the same performance bugs that always plagued data pipelines, now wrapped around an LLM. Rule 0: lock a quality baseline before you touch anything. Performance work is dangerous because the fastest version is often subtly less correct. Capture a baseline — representative dataset, current outputs, key quality metrics — and re-run it after every change. No optimization is accepted that regresses the baseline. Rule 1: profile, don’t guess. A pipeline matching 10,000 × 10,000 items “feels” model-bound, but the profile shows 90% of the time in a nested loop doing 100,000,000 comparisons. The model was never the problem. The usual suspects, in order of payoff: 1. The O n² match loop → index once into a hash/keyed join, then look up: ~100M comparisons become ~20k operations. 2. Per-record model/network round-trips → batch, or skip the model entirely for easy cases partition rows, is deterministic , rules for the easy ones, one batched model call for the hard ones . This is the same “don’t call the model” lever from cost, paying off twice. 3. Serial stages that could be parallel → bounded concurrency, sized to the real constraint provider rate limit, CPU, memory — not unbounded, which just moves the bottleneck and blows limits. 4. Recomputation → cache deterministic work embeddings, parsed inputs, reference lookups . After every change, re-run both axes — latency benchmark and eval vs baseline — and revert anything that regressed quality no matter how fast it is: change latency quality vs baseline hash-join was O n² 180s → 12s = baseline ✓ batch model calls 12s → 6s = baseline ✓ parallel stages ×8 6s → 1.8s = baseline ✓ The model is rarely the bottleneck — and “fast” should never be a guess about whether it’s still correct. You don’t have “a model.” You have a fleet — cheap and capable, primary and judge, this provider and that — and you need a layer that picks the right one, survives when one goes down, and can be stopped in seconds when it misbehaves. All three live behind the same gateway, and all three only work if the domain doesn’t care which model answered. Route by what the decision needs condition route to why ------------------------------- --------------------------------- ---------------------------------- is judge a different family from primary independent blind spots stakes == low / high volume cheap/small model most traffic; biggest cost lever ambiguous / high stakes capable model accuracy where it matters long-context / extraction-heavy the model measured best at it task fit only if you've measured php def route d - str: if d.is judge: return JUDGE MODEL independent from primary if d.stakes == "low": return CHEAP MODEL if d.needs long context: return LONG CTX MODEL return CAPABLE MODEL keep routing rules in ONE place the gateway , readable + testable — not per-call-site strings Fallback: survive a provider going down php def route chain req - list str : the ordered fallback chain if req.is judge: return JUDGE MODEL no silent fallback for a judge if req.stakes == "low": return CHEAP MODEL, CAPABLE MODEL fall UP on failure return CAPABLE MODEL, FALLBACK MODEL def complete with fallback req, deadline : for model in route chain req : if breaker model .is open : CHECK BEFORE calling — skip a known-dead provider continue remaining = deadline - now per-ATTEMPT budget vs an overall deadline if remaining <= 0: break try: resp = call model, req, timeout=remaining, idem key=req.idem key breaker model .record success return resp except Timeout, ProviderError : breaker model .record failure record on EVERY failure → the breaker can open return route to human req, reason="all models unavailable" degrade deliberately, don't error The non-obvious bits that make this correct: check the breaker before calling and record a failure on every caught error the common bug — doing both in the except means the breaker never proactively skips a dead provider . Use one overall deadline with per-attempt timeouts — the naive alternative, a fixed timeout reused per model, makes total latency N × timeout and blows the upstream budget. Idempotency is mandatory : a primary that timed out but actually completed and wrote a ledger entry must not be double-processed. And degrade deliberately — decide per capability whether a weaker fallback’s answer is acceptable or it should route to a human. One discipline ties routing and fallback together: eval every model on the path. A routed-to or fallen-back-to model is a different model, so potentially different quality. Your golden set should pass on the cheap model and the fallback model for the capabilities that use them — a fallback that quietly tanks quality is a worse outage than the one it covers. Record which model decided in the ledger , so outcome analysis can see whether the cheap or fallback model underperformed. The kill switch: an off that takes effect in seconds A deploy takes minutes you may not have. The switch must be runtime state every node checks: class Mode Enum : LIVE = "live" normal HUMAN ONLY = "human only" stop auto-execute; still propose to humans HALTED = "halted" stop deciding entirely def get mode switch, tenant id, capability - Mode: try: raw = switch.read f"killswitch:{tenant id}:{capability}" most specific or switch.read f"killswitch:{tenant id}" whole tenant or switch.read "killswitch:global" global flip lands everywhere return Mode raw if raw else Mode.LIVE except SwitchUnavailable: return Mode.HUMAN ONLY fail toward safe — NOT live, NOT halted a blip shouldn't self-DoS Four design choices make it trustworthy: fast propagation back it with shared state / pub-sub so a flip lands across all instances in seconds — a 30s local TTL is not “seconds” ; granular per-tenant and per-capability, so the blast radius of “off” matches the blast radius of the problem ; a middle gear HUMAN ONLY keeps proposing while stopping auto-execution — often you don’t need off , you need humans back in the loop ; and fail toward safe if a node can’t read the switch, assume the conservative mode . Degrade by design, and prove it Decide in advance how the system bends so failure isn’t a cascade. A circuit breaker stops you hammering a dead dependency — failing in milliseconds instead of behind 60s timeouts, which is exactly how a dependency outage becomes your thread-exhaustion outage. Under overload, shed load deliberately : slow down or route-to-human rather than crash. A system that degrades to “a human handles it” is still serving its purpose. These mechanisms are worthless if they only work in theory. Run chaos tests against agents: chaos test assert --------------------------- ----------------------------------------------------- kill the model provider decisions route to humans, not error out flip the kill switch auto-execution stops, fast, across all instances overload / burst sheds load / degrades, doesn't crash restart a node mid-decision in-flight decisions recover or fail safe idempotent dependency latency spike circuit breaker trips; no thread pileup If you haven’t exercised the kill switch and fallbacks under real failure, you don’t know they work — you’re hoping. The kill switch is the thing that lets you sleep. All of the above assumes a place to stand: a structure that lets you swap models without a refactor, an identity model that keeps each action attributable, and a right-sized set of infrastructure. Hexagonal architecture: keep the vendor out of your domain The model you ship on won’t be the model you started with. Providers leapfrog every few months; pricing changes; a region or compliance rule forces a switch. If your business logic is littered with vendor SDK imports and provider-shaped request objects, every one of those is a refactor. The fix is an old idea applied to a new problem: ports and adapters. Your domain depends only on ports — interfaces you define, in your terms. The messy outside model providers, datastores, queues lives in adapters that implement those ports. The domain never imports a vendor SDK. Define one gateway through which all model traffic flows, in your vocabulary: @dataclass frozen=True class ChatRequest: YOUR vocabulary — note what you deliberately DON'T expose messages: list dict schema: dict | None = None structured output your concept, not a provider's response format max tokens: int = 1024 no provider-specific knobs logit bias, etc. — they leak the vendor into the domain. class LLMGateway Protocol : shown sync for clarity; a real gateway is async/streaming def complete self, request: ChatRequest - ChatResponse: ... your types, not a provider's This is the same gateway that does cost measurement, rate limiting, routing, fallback, and the kill switch — every cross-cutting concern lives here once, not scattered. That’s not a coincidence: the chokepoint that makes the domain portable is the chokepoint that makes the system operable. Because the domain depends on a Protocol, tests inject a fake — no network, no spend, no flakiness — while adapters get their own integration tests against the real thing. Don’t rely on vigilance to keep the boundary clean; enforce it with a linter. An import-contract rule “the domain may not import vendor SDKs” turns the build red the moment someone slips a vendor import into domain code. Pair it with a naming convention — anything under adapters/ or suffixed per-provider may import that provider, nothing else may. Most ecosystems have an equivalent module-boundary lint in JS/TS, ArchUnit in Java, depguard in Go . It costs an extra translation layer and some upfront interface design; it buys one-file provider swaps, a testable domain, and a boundary the build maintains forever. Then "can we switch models?" stops being a project and becomes an afternoon. Identity: whose authority is each action carried with? An agent that acts on a user’s behalf has an identity problem a stateless API doesn’t. A request arrives as a user; the agent reasons, calls downstream services, maybe runs a while, maybe does work after the user has gone. At every hop: whose authority is this running with, and is it allowed to do this, for this user, in this tenant? Mishandle it and you’ve built the classic confused deputy — the agent using its own broad privileges to do something the user couldn’t. Two flows, two strategies. Short, synchronous < ~60s : propagate the user’s credential — the inbound token rides along to downstream calls, so every action runs with exactly the user’s authority safe only when every downstream does its own per-action authz . Long-running / deferred: you can’t hold the user’s token, so at the entry node mint a short-lived delegated grant — an asymmetrically-signed token scoped to this workflow, acting for the user, bound to one audience and one tenant, with least-privilege exact-match scopes, a revocable id, and a TTL in hours with a hard cap. The single most important check, on every hop , is that the identity’s tenant matches the request’s tenant: python def authorize token: str, request : AUTHENTICATE first: asymmetric signature downstream verifies but can't mint , reject alg=none, check exp/iat, aud == this service, jti not denylisted. Never trust a parsed identity. identity = verify grant token, expected aud=THIS SERVICE, leeway s=30 if identity.tenant id = request.tenant id: audit.violation "cross tenant", identity=identity, request=request raise Forbidden never proceed if request.action not in identity.scope: EXACT membership — no "write: " wildcard raise Forbidden f"out of scope: {request.action}" return identity This is the check that stops one tenant’s agent from ever touching another’s data — even through a bug or a prompt injection. Sign asymmetrically a shared HMAC secret lets any holder mint grants ; use least privilege and short TTLs a leaked short-lived grant is a small problem, a long-lived one is a breach ; validate at every hop injection happens after the edge ; and log cross-tenant denials as audit violations — they’re attack signal, never a silent 403. The point of all of it: every single thing your agents do is attributable, scoped, and bounded to the person and tenant it was meant for. Infrastructure: provision the spine, resist the cargo cult Standing up an agent platform swings between two failure modes: under-provisioning no audit store, no secrets management, one shared god-credential or over-provisioning a queue, three databases, and a service mesh for what is really a handful of stateless services . The actual shortlist: component why start with --------------- ----------------------------------------------------------- ------------------------------------ agent services the capabilities a few stateless services, autoscaled model gateway one chokepoint; only thing holding provider creds 1 audit datastore the append-only decision ledger canonical record a relational DB cache session/working state; backing for rate-limit + kill switch 1 e.g. Redis secrets store model keys, guardrail config, signing keys managed secrets manager object storage large inputs/artifacts + tiered regulated logs, encrypted 1 bucket Agent endpoints are structurally stateless HTTP services . If they propose decisions and don’t run long internal loops, you usually do not need a queue or event bus to start — sync request/response covers it. Notice what’s deliberately absent: a queue, a vector store, a second database. Add each only when a concrete workload demands it. Two disciplines matter more than the component list. Least-privilege IAM : don’t hand every service one broad credential — the gateway gets model:invoke and the model-keys secret; agent services get the audit store and their own object-store prefix but no model creds they call the gateway ; the audit DB user is scoped to its schema. Use per-service workload identities, never shared static keys, so a compromise’s blast radius is one component’s narrow permissions, not the platform. And encryption : a managed key for data at rest from day one, a dedicated signing key if the ledger is cryptographically signed, and per-tenant keys designed for even if you start single-tenant. The discipline isn’t adding exotic infrastructure — it’s not adding it, scoping every credential tightly, and putting the audit store and the gateway chokepoint in from day one. Operability is what turns a working LLM system into one you can actually run. Instrument the decision , not just the service — a metric catalog labeled by tenant and capability, tiered logs, a decision id threading metrics, logs, traces, and ledger, and alerts on symptoms with real thresholds. Funnel all model traffic through one gateway so cost is visible and bounded, then pull the levers in order: skip the model, batch, cache, right-size, trim — and profile before you optimize, baselining quality at every step. Treat the model as a routing decision , not a constant: cheap to cheap, hard to capable, an independent model to judge, failover so one provider's outage isn't yours — with an instant, granular kill switch and designed degradation you've actually proven under chaos. And stand all of it on a platform that keeps the vendor behind ports, threads tenant-checked identity through every hop, and provisions the spine without the cargo cult. The thread running through every section is the same: one chokepoint. The gateway that makes your domain portable is the gateway where you measure cost, route models, fail over, and flip the kill switch. Build that one seam well and operability stops being a scramble during incidents and becomes a property of the system. Series: Running LLM systems in production — Level 5 of 6: Operability.