{"slug": "nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router", "title": "Nvidia's latest solution to soaring enterprise AI costs is...a router?", "summary": "Nvidia unveiled NeMo Switchyard, a software router that directs AI prompts to different models to cut costs, claiming a 74 percent reduction in job completion costs versus using Claude Opus 4.8 alone, with a six-point accuracy tradeoff. The platform, announced alongside the open-weights Nemotron 3.5-30B-A3B-Lightning model, aims to make enterprise AI spending more manageable by routing simpler tasks to smaller, cheaper models.", "body_md": "Soaring AI infrastructure costs and model pricing, combined with uncertain returns on investment, threaten to stall enterprise adoption.\n\nTo make enterprise AI spend a bit more manageable, Nvidia this week unveiled a new software platform that blurs the line between expensive proprietary models and open weights alternatives.\n\nAnnounced alongside [Nemotron 3.5-30B-A3B-Lightning](https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/), Nvidia’s latest open weights model, NeMo [Switchyard](https://developer.nvidia.com/blog/route-ai-agent-workloads-across-models-with-nvidia-nemo-switchyard) is the GPU giant’s latest overture to enterprise. So what exactly is it? Well, it’s a router.\n\nThe idea is simple. Switchyard essentially functions as a proxy that sits between the inference server’s API endpoint and the models. But rather than sending every request to the same model, Switchyard can be configured to route prompts to different models in order to optimize for cost, latency, or output quality.\n\nBy routing some requests to smaller, cheaper, and potentially locally hosted AI models, Nvidia claims Switchyard can cut job completion costs by 74 percent relative to using Claude Opus 4.8 alone, albeit with an approximately six-point accuracy tradeoff.\n\n### The right tool for the job\n\nThe key metric in all of this is completion cost rather than price per token. A model might cost one-tenth as much as OpenAI’s or Anthropic’s top model, but if it requires 10x the tokens to complete the request, it isn't actually cheaper.\n\nCertain elements of an AI workload may benefit from a larger, smarter model, but not all do.\n\nFor example, it’d be overkill to ask Claude Opus to generate a title card or summarize a website. It’ll certainly work, but it’ll also cost a fortune compared to Haiku or a locally hosted model that’s been fine tuned just for that purpose. The fewer tokens you burn on the big smart model, the less expensive your API bill is going to be.\n\nNvidia software teams have spent the last several years developing models for this reason. The Lightning model announced this week is only its latest. The 30 billion-parameter MoE model is positioned as a low-latency, general purpose model that can either be used on its own or in conjunction with a larger, smarter model via a router like Switchyard.\n\nThe company has also developed several application-specific models. [Nemotron Parse](https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-v1.2) is one such example. “It’s a small model, one billion parameters, and it’s really good at one task, which is taking a PDF in and then explaining the context inside that PDF whether it’s charts or graphs or tables,” Joey Conway, senior director of AI software and models at Nvidia, explained in a recent interview with The Reg.\n\nMany frontier models struggle with this task because PDFs are designed by humans for humans, so by offloading that work to task-specific models, enterprises can not only improve the accuracy of their AI apps, but also reduce costs in the process.\n\nThis all might sound familiar: It's not the first time we’ve seen model routers employed as a cost-saving measure.\n\nBack when OpenAI launched GPT-5, ChatGPT would dynamically route prompts to different versions of the model based on their complexity.\n\nAs we wrote at the time, OpenAI’s router was likely [implemented](https://www.theregister.com/software/2025/08/13/openais-gpt-5-is-a-cost-cutting-exercise/1317079) to reduce the number of compute cycles spent on mundane tasks like rewording emails to sound more professional (\"not only … but also\").\n\nOpenAI wasn't alone in using routers to reduce model costs. The Wall Street Journal [recently reported](https://www.wsj.com/cio-journal/why-at-t-is-betting-big-on-open-weight-ai-a0ea03b1?st=t7aCLG&reflink=desktopwebshare_permalink) that AT&T has implemented a “smart router” of its own to automatically select which model to use. Switching from proprietary to open-weight models has reportedly saved the telecommunications giant between 80 and 90 percent in certain applications. Today about 25 percent of the company’s AI workloads are powered by open models. The company’s leadership expects that over the next few years that’ll climb to 70-80 percent.\n\n### The implementation challenge\n\nWhile the idea of offloading simpler requests to smaller, cheaper-running models sounds intuitive, it’s easier said than done. Title cards and web summaries are relatively straightforward to implement. Open source chatbots like Open WebUI have supported this kind of functionality for more than a year now because it just makes sense.\n\nHowever, sometimes it’s not obvious when and where these task models should be used. Switchyard is Nvidia’s latest attempt to simplify this by automatically routing requests to the right model for the job. However, it’s not the only approach Nvidia is exploring.\n\nAI agents and code assistants have the ability to work through problems and then generate skills — essentially standard operating procedures — documenting the process for future reference.\n\nThrough this iterative process, Conway suggests, agents could essentially teach themselves when and where they can get away with using a smaller, cheaper task model, and where a larger frontier model may be required.\n\n“We’re starting to see signs of this sort of agent and subagent type workflow,” Conway said, describing how a frontier model might function as an orchestrator that farms out work to smaller models that are faster and more specialized.\n\nIt reflects the way companies are structured, he said. “We have people who are specialists and then we have people who help orchestrate that and understand the complexity of the problem.”\n\nAs an added step, it’s possible for the agents to generate training data on the fly, which could then be used to fine-tune the models to operate more efficiently.\n\nRegardless of which approach ultimately wins out, anything that promotes enterprise AI adoption is a win for Nvidia. ®", "url": "https://wpnews.pro/news/nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router", "canonical_source": "https://www.theregister.com/ai-and-ml/2026/08/12/nvidias-latest-solution-for-soaring-enterprise-costs-nemo-switchyard-software-router/5286911", "published_at": "2026-08-12 19:00:10+00:00", "updated_at": "2026-08-12 22:18:52.708998+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products", "ai-tools", "large-language-models"], "entities": ["Nvidia", "NeMo Switchyard", "Nemotron 3.5-30B-A3B-Lightning", "Claude Opus 4.8", "OpenAI", "Anthropic", "AT&T", "Joey Conway"], "alternates": {"html": "https://wpnews.pro/news/nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router", "markdown": "https://wpnews.pro/news/nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router.md", "text": "https://wpnews.pro/news/nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router.txt", "jsonld": "https://wpnews.pro/news/nvidia-s-latest-solution-to-soaring-enterprise-ai-costs-is-a-router.jsonld"}}