cd /news/artificial-intelligence/nvidia-unveils-nemotron-3-5-lightnin… · home topics artificial-intelligence article
[ARTICLE · art-92026] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Nvidia unveils Nemotron 3.5 Lightning and NeMo Switchyard to slash enterprise AI costs

Nvidia released Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model, and NeMo Switchyard, an open-source Rust library for routing AI workflow steps to the most appropriate model, on August 11, 2026. The model activates only about 3 billion parameters at a time, making it cost-efficient for high-volume agentic tasks, while Switchyard optimizes cost, latency, and capability across models.

read2 min views2 publishedAug 11, 2026
Nvidia unveils Nemotron 3.5 Lightning and NeMo Switchyard to slash enterprise AI costs
Image: Cryptobriefing (auto-discovered)

Via dwglogo.com

The new 30-billion-parameter model and open-source routing library let companies stop choosing between smart AI and affordable AI

Nvidia released Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model, alongside NeMo Switchyard, an open-source library that directs each step of an AI workflow to the most appropriate model available.

What Nemotron 3.5 Lightning actually does #

The model uses a mixture-of-experts architecture, which means that while it has 30 billion total parameters, only about 3 billion are active at any given moment. This design makes Nemotron 3.5 Lightning particularly suited for high-volume, specialized agent tasks. The kinds of operations that enterprises run thousands of times per day, like document parsing, data extraction, or customer query classification, don’t need frontier-scale reasoning.

Nemotron 3.5 Lightning extends a model family that has been growing steadily since late 2025, when Nvidia rolled out variants including Nano, Super, and Ultra. Each targets a different slice of the performance-cost spectrum, and Lightning slots in as the option optimized for agentic workloads that need to run cheaply at massive scale.

NeMo Switchyard: the traffic controller #

NeMo Switchyard is an open-source Rust library, now available on GitHub, that handles LLM traffic routing and API translations across different models. When an AI agent kicks off a multi-step workflow, Switchyard evaluates each step and routes it to whichever model fits best based on parameters like cost, latency, or capability requirements. A simple classification task might go to Lightning. A complex reasoning step might get routed to a larger frontier model.

The release of Nemotron 3.5 Lightning and NeMo Switchyard was made public on August 11, 2026.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nvidia-unveils-nemot…] indexed:0 read:2min 2026-08-11 ·