{"slug": "nvidia-nemotron-3-5-lightning-just-landed-in-cline", "title": "NVIDIA Nemotron 3.5 Lightning just landed in Cline", "summary": "NVIDIA released Nemotron 3.5 Lightning, a 30B mixture-of-experts model with 3B active parameters, now available in Cline, an AI coding assistant. The model, distilled from NVIDIA's Nemotron 3 Ultra, supports up to 1M token context and offers up to 4x higher throughput than similar-sized open models, scoring 86.2 on PinchBench, 54.3 on SWE-Bench Verified, and 73 on AA-Omniscience non-hallucination. It is designed for high-volume agentic coding workflows, and Cline users can access it via OpenRouter.", "body_md": "[Models](https://cline.ghost.io/tag/models/)\n\n# NVIDIA Nemotron 3.5 Lightning just landed in Cline\n\nNVIDIA Nemotron 3.5 Lightning is now available in Cline, bringing a 30B MoE model with 3B active parameters, up to 1M context, and high throughput for agentic coding workflows.\n\nNVIDIA just released Nemotron 3.5 Lightning, a new customizable open model built for always-on agents and it is now available in Cline.\n\nNemotron 3.5 Lightning is a 30B MOE model with 3B active parameters, distilled from NVIDIA’s frontier Nemotron 3 Ultra. It supports up to 1M token context window and is designed for always-on agents that need to move through a lot of model calls quickly. If you're using Cline, this matters.\n\n## One task, a lot of model calls\n\nSome parts of an agentic coding task need deeper reasoning. Others are much more execution heavy like finding the right file, reading an implementation, running tests, checking an error, making the next edit.\n\nBut every one of those steps still calls a model.\n\nNVIDIA's approach with Nemotron 3.5 Lightning is to make those repeated calls fast while maintaining the agentic capability needed for specialized tasks. The model is trained for popular agent harnesses and designed for high throughput tasks. .\n\nThat makes it an interesting fit for executing heavy Cline tasks where the agent may go through the read, edit, test, and retry loop many times before the work is done.\n\nIn Cline, speed compounds over the course of a task. One faster response is nice, but faster responses across dozens or hundreds of turns can change how quickly the whole agent loop moves.\n\n## A 30B model with 3B active parameters\n\nNemotron 3.5 Lightning uses a hybrid mixture-of-experts architecture with 30B total parameters and 3B active parameters during generation. It’s distilled from NVIDIA’s frontier Nemotron 3 Ultra model and supports a context window of up to 1M tokens.\n\nNVIDIA's preliminary throughput numbers are where Lightning stands out.\n\nDespite the smaller active footprint, Nemotron 3.5 Lightning scores 86.2 on PinchBench, 54.3 on SWE-Bench Verified, and 73 on AA-Omniscience non-hallucination. Compared to other leading open models of similar size, Nemotron 3.5 Lightning offers up to 4x higher throughput – placing it on the accuracy-speed Pareto frontier for high-volume agent workloads.\n\nIt is an open, customizable model that can be post trained for specialized workflows and deployed locally, at the edge, in the datacenter, or in the cloud.\n\n## Using Nemotron 3.5 Lightning with Cline\n\n### Prerequisites:\n\n- A\n(free to create)__Cline account__ - Cline installed in your IDE, or the Cline CLI\n\n### Setup\n\n- Open Cline\n- Go to Settings\n- Select OpenRouter as your API provider\n- Select\n`nemotron-3.5-lightning-30b-a3b`\n\nfrom the model dropdown - Done!\n\nNemotron 3.5 Lightning is an interesting addition to the growing set of open models built around agentic workloads. Not every step in an agent loop needs the same kind of model. Lightning focuses on making the repeated calls inside that loop fast enough to keep the whole task moving.\n\nTry Nemotron 3.5 Lightning in Cline and let us know what kinds of tasks it handles well.\n\nFor questions or to share what Lightning handles well, join the conversation on[ Discord](https://discord.gg/cline?ref=cline.ghost.io) &\n\n[.](https://www.reddit.com/r/CLine/?ref=cline.ghost.io)__Reddit__", "url": "https://wpnews.pro/news/nvidia-nemotron-3-5-lightning-just-landed-in-cline", "canonical_source": "https://cline.ghost.io/nvidia-nemotron-3-5-lightning-available-in-cline/", "published_at": "2026-08-11 14:36:23+00:00", "updated_at": "2026-08-11 14:49:12.353870+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-tools"], "entities": ["NVIDIA", "Nemotron 3.5 Lightning", "Cline", "OpenRouter", "Nemotron 3 Ultra"], "alternates": {"html": "https://wpnews.pro/news/nvidia-nemotron-3-5-lightning-just-landed-in-cline", "markdown": "https://wpnews.pro/news/nvidia-nemotron-3-5-lightning-just-landed-in-cline.md", "text": "https://wpnews.pro/news/nvidia-nemotron-3-5-lightning-just-landed-in-cline.txt", "jsonld": "https://wpnews.pro/news/nvidia-nemotron-3-5-lightning-just-landed-in-cline.jsonld"}}