cd /news/artificial-intelligence/snowflake-adds-dynamic-model-routing… · home topics artificial-intelligence article
[ARTICLE · art-101318] src=cio.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Snowflake adds dynamic model routing to Cortex AI Gateway to cut enterprise AI costs

Snowflake introduced dynamic model routing for its Cortex AI Gateway, which automatically directs workloads to the most cost-effective model based on cost, performance, and latency requirements. The company's internal tests show token efficiency improvements of up to three times for certain tasks, but analysts caution that token efficiency does not equal cost savings and that operational complexity shifts to the governance layer.

read5 min views4 publishedAug 18, 2026

Snowflake on Tuesday unveiled a dynamic model routing capability for its Cortex AI Gateway, designed to help enterprises reduce AI spending by automatically directing workloads to the most appropriate model based on cost, performance, and latency requirements.

The new capability, which is expected to be in private preview soon, will allow enterprises to define which models they approve for use and the tradeoffs they want the system to prioritize, such as cost, performance, and latency, for an individual application or workload, CEO Sridhar Ramaswamy wrote in a blog post.

Once those policies are defined, Cortex AI Gateway then evaluates each task against those policies and real-world model performance and cost data to determine which model should handle the workload in the most efficient manner, Ramaswamy added.

Further, the CEO pointed out that Cortex AI Gateway also creates a feedback loop by evaluating the quality of a model’s output after it completes a task, which allows the routing system to adjust its decisions as model capabilities, pricing, and performance change, with the aim of continuously optimizing the balance between quality, cost, and latency.

According to Snowflake’s internal benchmarks, the new capability can improve token efficiency compared with using a frontier model for every task.

In one internal test, agents using dynamic routing built a dbt pipeline with up to three times greater token efficiency than a frontier-model-only approach while maintaining the same quality, the company said in a statement. In another test, engineering teams completed the same number of pull requests with 25% greater token efficiency, it added.

The new capability will have the largest impact on high-volume, low-complexity workloads where many requests do not require frontier-model reasoning, like classification, extraction, summarization, routine data engineering, and repetitive agent steps, said Stephanie Walter, practice lead of AI stack at HyperFRAME Research. “Routing those requests to smaller models could materially reduce inference costs while preserving expensive models for genuinely difficult tasks,” said Stephanie Walter, practice lead of AI stack at HyperFRAME Research.

Agentic applications could specifically benefit from dynamic model routing, said Advait Patel, senior site reliability engineer (SRE) at Broadcom.

“An agent loop spends most of its steps on plumbing, reading a file, parsing a result, and picking the next call. Very few of those need deep reasoning, but they all hit the same model today. When I pulled telemetry on our own coding agent usage, the spend wasn’t in the hard problems at all. It was the volume of ordinary calls,” Patel said.

However, Walter cautioned that enterprises should not treat token efficiency as the same as cost savings, especially in agentic applications, despite Snowflake’s “promising” internal benchmarks.

“Enterprises must also measure retries, failed tasks, latency, human correction, and the cost of operating the routing layer,” Walter noted.

More so because routing, despite removing the repetitive model-selection work from individual applications, shifts operational complexity into the orchestration and governance layer and doesn’t eliminate it completely, according to Phil Fersht, CEO of HFS Research.

“Enterprises would still need to determine which models are approved, establish routing policies, monitor quality, control costs, and manage security and compliance,” Fersht said, adding that if policies are not defined well, the system can make a poor decision, which at scale, could either produce inconsistent outcomes or unnecessary costs.

That shift of operational complexity into the governance layer, according to Manoj Chandra Jha, principal analyst at Nord-IQ Research, could be challenging for most enterprises: “Short-term complexity can rise, since most teams lack the governance and monitoring maturity routing now requires.”

The governance burden also has implications for developers, who will have to account for routing decisions as another variable when building and troubleshooting applications.

“Dynamic routing makes visibility essential. If different requests go to different models, developers need to know which model handled a request, why it was selected, and whether the result met expected quality and performance levels,” said Robert Kramer, managing partner at KramerERP.

“When something breaks, they need to determine quickly whether the fault came from the application, the model, or the routing decision. That third failure mode is new, and it is the one teams are least equipped to diagnose today,” Kramer added.

Snowflake, however, is looking to address concerns around changes in routing decisions driven by model pricing changes.

It would integrate Cortex AI Gateway with its AI coding assistant CoCo’s existing role-based access and tagging framework, which will allow enterprise administrators to set default models, attribute AI usage to teams or cost centers, establish per-user quotas, and receive alerts as consumption approaches predefined limits.

These controls could help enterprises maintain visibility into how routing decisions affect AI spending as models, pricing, and workloads change, the company said.

That enterprise focus on controlling AI spending via model selection and routing hasn’t escaped the attention of other vendors.

Nvidia has been expanding its efforts around model routing, while Cloudflare and OpenRouter have also emerged as players in the space, reflecting growing interest in helping enterprises route workloads across multiple models based on factors such as cost, performance, and capability.

The shift, according to Fersht, is partly a consequence of the growing number of models available to enterprises and the differences between them in cost, performance, latency, and capabilities.

That shifts the strategic value towards the layer that decides which model to use and orchestrates it across enterprise workflows, Fersht noted.

However, Patel cautioned that enterprises should evaluate model routers based on the level of control and transparency they provide.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @snowflake 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/snowflake-adds-dynam…] indexed:0 read:5min 2026-08-18 ·