An LLM gateway is a small system with a big trust load. It's small enough for one on-call to own, and it holds your org's model credentials, its spend attribution and its audit trail. That combination shapes how you run it: conservative where the trust is, pragmatic everywhere else.
The runbook set I keep for one has four incidents, not forty. These are the ones that actually happen, and every runbook has the same four steps:
The last step is the one teams skip. An incident without its artifact is a memory, and the memory is the next incident.
Before any runbook, the pager needs rules, or every incident review starts with "we ignored the alert". Three severities with fixed meanings:
Two rules keep the pager trusted:
crit escalates once, after 4 hours, then goes quiet. An alert that pages forever gets muted, and a muted Here's how my default thresholds split. Treat them as starting points to tune, not laws:
| Signal | warn (this week) | crit (now) |
|---|---|---|
| Budget, per team or feature | 95% of the hard limit | 100% (hard) and 110% (overage) |
| Cost anomaly | a feature's 7-day cost ≥ 3× its 30-day average | a team's daily cost ≥ 5× its 30-day daily average, or one request over your per-request cap |
| Metering completeness | below the 99.9% SLO | below 99.5% for an hour |
| Provider breaker | open, and fail-over worked | open, and the alias has no secondary (users get 503s) |
| Metering bus lag (p95) | over 5 minutes | no page: dashboards must say "stale" instead |
| Credential age | past its rotation schedule | suspected leak: revoke first, investigate second |
Now the four incidents.
Detect. The provider's circuit breaker opens, or your services start reporting 503 model_unavailable on an alias. The provider's status page is corroboration, not detection. Your breaker sees it first.
Assess. One query: request outcomes by provider over the last 15 minutes. One provider failing is a provider incident. Every provider failing at once is your gateway, and that's a different runbook.
Act. If the alias has a secondary binding, fail-over already happened. Your job is to confirm it: watch the share of responses marked X-LLM-Router: failed-over=1 (or metering events where routed_to != requested_binding), and check the cost delta of the secondary model. If the alias has no secondary, bind one in the model registry. That's a registry change, not a deploy, which is the whole point of routing through aliases.
Close. The record names the window, the fail-over share, the cost delta, and the provider's acknowledgment if one came. Add a note to that provider's availability history. Your provider mix is a portfolio, and you rebalance it on that evidence, not on loyalty.
Detect. The crit cost alert, or finance asking. Finance's question is the same alert with a slower detector.
Assess. Look at the per-feature, per-user and per-model views, in that order:
Act. It depends on what you found:
Close. Name the cause class (loop, price, fail-over, one-off), the cost, and the budget state afterwards. If a budget should change, that goes through the normal budget workflow, not the incident. The incident doesn't grant budgets; it finds them.
Detect. A wall of 401/ 403, a degraded identity platform, or "my token works in staging but not in prod."
Assess. This is why denial reasons deserve their own label instead of one unauthorized bucket. Split the denials by reason and the shape names the cause:
token_invalid spikes.token_expired across the whole org.scope_missing or role_denied in one team.
Act. For the JWKS shape, force a JWKS refresh through the gateway's admin path and check the identity platform's key state. For the TTL shape, the fix is upstream in the token mint; your job meanwhile is checking whether the affected services can degrade safely. For the grant shape, restore the grant through the normal access workflow, so the fix carries its own audit trail.
Close. Record the shape, the blast radius (which teams, which services), the recovery action, and whether anyone used break-glass. If they did, the break-glass audit line goes in the record. This is the incident that touches the trust load, so the record matters most here.
Detect. Metering completeness drops below the 99.9% SLO, the monthly reconciliation finds a gap, or bus lag alerts. Below 99.5% for an hour is an incident, because your books are now wrong.
Assess. Find where the events went missing. Missing from the bus is a transport loss. On the bus but missing from the rollups is an aggregation loss. Bus lag catches the first; completeness catches the sum. If lag is normal while completeness drops, look at the rollups.
Act. For a transport loss, follow the bus's own runbook while the gateway's bounded queue holds events, and watch completeness recover. For a rollup loss, recompute the rollups from the bus. The recompute is a job, and its run is itself an audit event.
One rule I don't bend: never back-fill metering from the provider's invoice. The invoice is the reconciliation's check, not metering's source. Back-fill from it and your books become the provider's books, with no team, feature or user attribution, and next month's reconciliation has two gaps instead of one.
Close. Record the window, the lowest completeness during it, the recompute if one ran, and the expected variance for month-end, so the finance close isn't a surprise.
The cheapest on-call improvement I know: write each runbook's assess query into the alert itself, so whoever gets paged at 3 a.m. doesn't have to remember it. A Prometheus sketch with the metric names used above. Thresholds come from the table; the for: durations are illustrative.
groups:
- name: llm-gateway
rules:
- alert: LLMProviderBreakerOpen
expr: max by (provider) (llm_provider_breaker_state) == 1
for: 2m
labels: { severity: warn } # raise to crit if the alias has no secondary
annotations:
first_query: 'sum by (provider, outcome) (increase(llm_requests_total[15m]))'
- alert: LLMMeteringBooksWrong
expr: llm_metering_completeness < 0.995
for: 1h
labels: { severity: crit }
annotations:
first_query: 'max_over_time(llm_metering_bus_lag_seconds[1h])' # high = transport, normal = rollups
- alert: LLMMeteringBelowSLO
expr: llm_metering_completeness < 0.999
for: 1h
labels: { severity: warn }
The cost alerts deliberately live in the budget service, not in Prometheus. They come from rollups, they need the state-entry dedup, and they should carry the budget percent, the forecast and the top features in the webhook payload, so the person who gets it can act without opening five dashboards.
Credential drift. Every credential the gateway holds rotates on a schedule (90 days is my default for provider credentials) through a two-secret swap: mint the new version, verify it with a tiny tagged probe call, flip the binding, and only then retire the old one. On a suspected leak, the order flips: revoke first, investigate second. The schedule is also what bounds a leak's exposure window: a 90-day rotation means a 90-day exposure, not a forever one.
The degrade-mode direct keys. If the gateway is down, some services that can't afford the gap may keep a scoped, audited direct-provider key. Fine, but every such key needs a named owner and a deletion review date. A direct key without an owner is a static key with a runbook around it.
401/ 403 wall; first query is denials by reason; close with the shape, blast radius and any break-glass use.
This is a condensed version of Chapter 7 of my book The AI Gateway Playbook, which also covers the audit-log retention contract, keyless credential rotation, the metrics catalog behind these alerts, and the deployment and migration checklists. Related reading: MCP at 90+ tools: the catalog breaks before the server does.
Disclosure: this article and the book were drafted with AI assistance from my own experience designing and running this kind of platform. All examples are generic.