# Cloudflare Launches Auto Router to Cut AI Costs by 30%

> Source: <https://insideai.news/news/ai-in-business/cloudflare-auto-router/13326/>
> Published: 2026-09-30 15:06:55+00:00

**September 30, 2026, (Inside AI)** — Cloudflare has launched the public beta of its Auto Router, a new feature within AI Gateway that automatically routes each AI request to the most cost-effective model capable of handling the task. Early internal tests show up to **30% cost savings** compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.

The Auto Router aims to solve a growing problem for enterprises: as AI adoption matures, organizations struggle to control token spend without hindering employee productivity. According to Cloudflare, the best savings are the ones users never notice. The router intelligently selects models based on task complexity, ensuring that simple tasks like summarizing an email do not consume expensive frontier model resources, while complex coding or security tasks still get access to top-tier models when needed.

Cloudflare's own experience tracking AI spend revealed that managing costs requires a multipronged approach. Previously, the company discussed setting budgets and limits, and linking employees to their AI usage. However, individual users still manually select models in many harnesses, often defaulting to overkill models for routine work. The Auto Router addresses this by making intelligent decisions on the user's behalf, reducing spend automatically while preserving access to capable models when necessary.

In internal evaluations using Cloudflare's OpenCode deployment and Cloudflare OS, the Auto Router delivered performance comparable to frontier models for coding tasks. On a general knowledge work benchmark with **97 tasks** and three samples per model per task, the router achieved similar performance to state-of-the-art daily-driver models at **80% the cost of Sol** and **35% the cost of Opus**. The benchmark simulated common workflows across email, calendars, Slack, files, travel, and finance.

**Read:** **Microsoft Copilot Super App**

One key insight from Cloudflare's research is that lower token prices do not always produce lower-cost outcomes. A model that appears cheaper per million tokens may use disproportionately more tokens to solve a problem. Therefore, the router minimizes predicted trajectory cost, not just load-balancing by dollars per million tokens. This approach accounts for the "jagged frontier" across models, where the ability to solve a problem often exists somewhere in a portfolio of models. The router's job is to choose the right model for each task while balancing quality and price.

Technically, when a request is sent to the Auto Router, AI Gateway first builds a pool of models that can serve it, filtering out those that do not support the request format or execution mode, and considering credentials, billing configuration, access control policies, and spend limits. It also filters out unhealthy upstream providers during downtime. For remaining candidates, the router analyzes a compact view of the conversation, prioritizing recent messages, and sends it to a multi-head classification model running on Workers AI and deployed on GPUs across Cloudflare's edge network.

The classifier assigns probabilities across **14 task categories** (like coding, planning, research, data analysis) and rates the request on four dimensions: complexity, ambiguity, stakes, and dependence on earlier context. A separate scoring matrix combines these signals with model benchmark results to estimate how well each model fits the request. The router then combines expected quality with each model's input and output token prices. On straightforward requests, price carries more weight, allowing smaller models to win when capable. As difficulty rises, the cost penalty falls, giving stronger models more room to win.

For long agentic sessions like debugging or coding, cost is less driven by the model's list price than by the cost of cache reads, which grows with session length. Switching models throws the cache away and forces a new model to write the whole context again. The Auto Router accounts for this by applying a switching penalty that grows with the number of tokens already in context. Within a turn, the cache is hot, so switching rarely pays off. Across turns, a model that still holds a live cache for the session is priced at its cheaper cache-read rate, while every other candidate is priced at the full cost of rewriting the context. This means the deeper the conversation, the more a switch has to earn back through higher quality results that use fewer tokens overall or cheaper cache rereads.

The two-stage architecture (task and dimensions classifier to scoring matrix) makes routing decisions legible, as each task's predicted category and complexity can be inspected to see how it translated into the model choice. Adjusting the router when a new model is released does not require retraining; only its benchmark-derived weights are added to the scoring matrix. The same classifier can support different routing profiles, and Cloudflare plans to release other routers in the future, including one that selects the highest expected quality without applying the cost tradeoff.

**Read:** **Context Engineering Overtakes Prompting**

Looking ahead, Cloudflare intends to expand the models offered through the Auto Router, include zero-data-retention requirements when filtering models, account for provider capacity when selecting models, select the appropriate reasoning or thinking level for each request, add full support for the Responses API and WebSockets, and explore structured decision models as a first-pass classifier. The Auto Router is free while in beta.
