# Nvidia Drives Bang For The Buck With New GenAI Model And Router

> Source: <https://www.nextplatform.com/ai/2026/08/11/nvidia-drives-bang-for-the-buck-with-new-genai-model-and-router/5286365>
> Published: 2026-08-11 16:07:35+00:00

# Nvidia Drives Bang For The Buck With New GenAI Model And Router

Nvidia is continuing to grow its reach and influence in AI in its efforts to remain the indispensable player in the ever-expanding and highly lucrative field. Today’s focus is on the technology itself, with the vendor adding to its [portfolio of open AI models](https://www.nextplatform.com/code/2026/03/18/the-open-agentic-ai-world-according-to-nvidia/5209529) and adding an open source model library that organizations can use to find the right model for every step in an agentic AI workflow, rather than relying on a default model or having to sort through their options on their own.

However, even as the world’s wealthiest company continues to build out its architectural offerings to help drive AI operations and adoption, it also is branching out into areas of increasing concern to the AI industry, which in recent weeks has meant security and [datacenter funding](https://www.nextplatform.com/ai/2026/07/21/genai-hardware-investments-are-way-ahead-of-model-and-platform-revenues/5275842).

Look at any layer of the AI industry, and more likely than not, you are going to see Nvidia in there.

In the model space, Nvidia is expanding its growing list of open models that include a range of Nemotron 3 models. Nemotron 3.5 Lightning falls in line with the rest of the Nvidia’s Nemotron models, which are designed to operate as part of what it calls systems of models, or model ensembles, with different models used for different steps in the workflow. Systems of models is not precisely the same as a mixture of experts, or MoE, which is an approach to building modern GenAI models that has dozens to hundreds of distinct experts activated selectively within a particular model. It is another layer of abstraction upwards.

“Efficient agents need a system of models, not just one model,” Kari Briski, vice president of generative AI at Nvidia, said in a prebriefing ahead of the launch. “Agent solutions and workflows make many model calls. After a plan is devised, it is carried out in steps, or what we call turns. Different steps require many tasks at different levels of intelligence.”

Lightning also comes at a time when [open models continue to roll into the market](https://www.nextplatform.com/ai/2026/07/21/genai-hardware-investments-are-way-ahead-of-model-and-platform-revenues/5275842), particularly from vendors from China, such as Moonshot AI’s Kimi K3, [DeepSeek](https://www.nextplatform.com/ai/2025/01/27/how-did-deepseek-train-its-ai-model-on-a-lot-less-and-crippled-hardware/1656116)’s various models, and Alibaba’s Qwen family. Stateside, Meta Glimmer this week rolled out Muse Glimmer, the first in a promised set of open models. (Meaning open weights, but not necessarily open source code for the training algorithm for the models.)

Briski noted that over the past year, Nvidia has rolled out its Nemotron Nano, Super, and Ultra, models, with each made for particular operations. Nano, which was released in December 2025, is a high-throughput, efficient small language model designed with [real-time AI agents](https://www.nextplatform.com/ai/2026/04/27/the-genai-battle-shifts-from-frontier-models-to-agentic-platforms/5218897), multi-step tool calls, advanced math, and coding. Nemotron 3 Super is a 120-billion parameter open-weight model with an MoE design for efficient agentic AI workflows, complex reasoning jobs, and accurate tool calling in enterprises.

### Adding To The Nemotron 3 Family

Nemotron 3 Ultra is a reasoning model that weighs in at 550-billion parameters and is used for high-throughput and long-running agents running multi-step workflows like agent orchestration, deep repository coding, research relying on multiple sources, and enterprise automation.

“Lightning brings together the best of our family,” Briski said. “The knowledge of Ultra packed into the size of our Nano. And Lightning is built for always-on agents handling a constant stream of specialized tasks, a supercharged workhorse model for high-volume AI that is fast, accurate, and efficient and delivers four times the throughput in its class at the same intelligence index as our super model, open for developers to post-train and optimized for specialized tasks.”

It also can run on Nvidia’s Jetson, GeForce RTX, DGX Spark, and DGX Station systems, allowing developers and their agents to use their own hardware to run operations instead of having to call a cloud API for every job.

The 4X output speed results in a 30 percent speed up in completing agentic AI tasks over other models, according to Nvidia. Organizations can post-train the model via its NeMo platform on their own data and tools for specialized tasks. Nvidia also is publishing its training data and techniques and is releasing Nemotron-RL-Agentic-Terminal-Pivot, an open dataset for training coding agents.

Briski said results from PinchBench testing – which rates open models on such agentic tasks as coding, research, and file management – illustrate that Lightning is up to 30 percent faster than comparable Qwen models, and faster than other frontier AI models, like Google DeepMind’s Gemma. Check it out:

The speed is important, she said, because “always-on” agents make thousands of models calls and latency matters. Lightning leverages Nvidia’s hybrid Mamba transformer architecture and multi-token prediction technology. According to the company, Lightning outperformed other open and proprietary models, as shown in the results below involving CodeRabbit, Harvey, and Lila in various tasks, and performed well against Nemotron 3 Super for CrowdStrike.

Take a look:

Along with Lightning, Nvidia also released the NeMo Switchyard routing library for AI models to allow organizations to choose the right model for the right task in the agentic workflow.

“The best model changes as the workflow evolves,” Briski said. “An agent has many states when completing tasks. The agent state changes as the tools return results, errors occur, or some tasks become routine. The router continuously balances quality, latency, and cost. A fixed model choice can't adapt to those changes. Routing isn't a new concept. We already rely on it everywhere from Internet traffic and phone networks. The challenge isn't just having many models, it's choosing the right one at each step of the workflow.”

AI model routers aren’t new. There are a range of vendors, from OpenRouter and Not Diamond to LiteLLM, and they’re trying to solve the same problems. Always-on agents or those running long, autonomous jobs can call on models and wrack up high tokens costs, driving up the expenses of running AI workloads and risking accuracy. Microsoft in May [unveiled Switchcraft](https://www.microsoft.com/en-us/research/publication/switchcraft-ai-model-router-for-agentic-tool-calling/), which the vendor says increases accuracy by almost 83 percent and reduces inference costs by 84 percent.

Nvidia is looking to offer its open model lineup with Switchyard as a way of giving developers a single place to get both the AI workflow and routing capabilities.

“Switchyard gives developers and platforms a way to define their own model pools, routing criteria, and policies,” Briski said. “It routes frontier models to reasoning intensive steps and lightning. To execution tasks where speed and efficiency matters. Now that is the power of a system of models, matching the right model to each step of the workflow.”

She said the chart below shows how Switchyard delivers better accuracy at a third of the cost of Anthropic’s Opus 4.8 model alone, and with improvements in both accuracy and cost when combining Opus 4.8 with Lightning and Gemma 3.

All of this comes as Nvidia is moving deeper into other areas of generative and agentic AI, including security. Soon after reports of OpenAI’s GPT-5.6-Sol and another unreleased frontier model autonomously breach Hugging Face, Nvidia led the development of a consortium of more than three dozen companies launched the Open Secure AI Alliance, arguing that protections used to defend infrastructure from AI threats need to be built on open models and tools that all defenders can adapt and deploy. The consortium now includes more than 120 members and is taking input as it develops new guidelines to enhance agentic AI cybersecurity.

In addition, the vendor reportedly is [staffing up an AI safety and security team](https://www.businessinsider.com/nvidia-staffs-new-ai-safety-team-push-for-open-models-2026-8), making another move to protect open models. The company is advertising for a security research engineer

To help organizations pay for their AI infrastructure, Nvidia is teaming with global investments firms Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to create a $500 billion fund. Enterprises, smaller companies, startups, and government agencies are looking for ways to build out their AI infrastructures.

The fund will help, according to Nvidia co-founder and chief executive officer Jensen Huang.

“In AI, compute is revenue,” Huang [said in a statement](https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital), pointing to the vendor’s DSX AI factories and the broad adoption of other Nvidia technologies, which are “supported by a deep global ecosystem of developers, customers and offtakers. That is why we are bringing the world’s leading long-term capital providers together to independently underwrite AI infrastructure.”
