Routing LLM requests by cost and latency means sending each request to the cheapest or fastest model that still meets the quality bar, rather than hardcoding a single model for everything. Production traffic isn't uniform: routine lookups, complex troubleshooting, and background jobs have different cost and latency requirements, even within a single app.
Using an inference router like DigitalOcean Inference Router involves defining request categories, assigning each a selection policy (cost, latency, or a benchmarked best fit), and setting fallback models so an unavailable model doesn't break the request. Cost is the price per token (which can vary by up to 100x across a model catalog); latency is mainly the time to first token (TTFT). The two don't move together, so route per task rather than using a single blended score. Key takeaways:
| Workload | Priority | Policy |
|---|---|---|
| Real-time chat | Latency | Speed-optimized (TTFT) |
| Routine lookups | Cost | Cost-optimized | | Complex reasoning | Quality | Benchmarked "optimal" | | Batch/background | Cost | Cost-optimized |
Fallback models keep requests completing when the preferred model is down or rate-limited.
Cache-aware routing matters too: switching models to save a fraction of a cent can break a cached prompt and cost more overall.
Inference Router tasks pair a model pool with a selection policy: preset tasks default to a benchmarked Optimal policy; custom tasks choose Cost Efficiency, Speed Optimization (TTFT), or Manual Ranking.
Fallback models handle unmatched traffic. X-Model-Affinity preserves cache reuse, and X-Routing-Max-Switch-Spend-Pct (default 20%) caps the cost of switching mid-session.
It's a drop-in change. Set "model": "router:your-router-name". Routing decisions typically resolve in about 200ms, billed at the serving model's standard rate with no separate router charge during public preview.
What is LLM routing?
LLM routing is the practice of directing each request to the model best suited to it—by cost, latency, or task fit—instead of sending every request through one model regardless of what it needs. The DigitalOcean Inference Router implements this as a managed feature, so teams get task-aware routing and fallback handling without building it themselves.
How do you model cost-per-token for LLM inference? Multiply input tokens by the model's input rate and output tokens by its output rate, then sum the two, since the rates are usually very different. Because rates vary widely across a model catalog, this is best tracked per task rather than as one blended number. The DigitalOcean Inference Router reports token usage and cost-relevant metrics per model and per task in its Analyze view.
What metrics matter for LLM inference observability?
Time to first token (TTFT), time per output token (TPOT), and inter-token latency (ITL) are the core metrics for judging responsiveness. A latency-optimized routing policy should be measured against these, not just overall request time. The DigitalOcean Inference Router uses TTFT specifically as the basis for its Speed Optimization selection policy.
What causes cold start latency in GPU inference, and how do you avoid it?
Cold starts happen when a model has to load onto available GPU capacity before it can serve a request, rather than running on an already-warm instance. Pooled serverless capacity and routing policies that keep related traffic on one model reduce how often this happens. The DigitalOcean Inference Router uses model affinity to keep a session's requests on the same warm model, rather than triggering repeated cold starts across models.
Can I route by both cost and latency at the same time?
Yes—typically by defining separate tasks or policies for different request types rather than one blended score. For example, cost-first for background work and latency-first for real-time chat. The DigitalOcean Inference Router supports this directly. Each task in a router can have its own model pool and its own Cost Efficiency, Speed Optimization, or Manual Ranking policy.