cd /news/artificial-intelligence/nemotron-3-5-lightning · home topics artificial-intelligence article
[ARTICLE · art-93619] src=tokenstead.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Nemotron 3.5 Lightning

NVIDIA released Nemotron 3.5 Lightning on 2026-08-11, a 31.6B-parameter mixture-of-experts model with ~3.6B active parameters per token, hybrid Mamba-Transformer architecture, multi-token prediction, and a 1M-token context window, available as open weights under the Linux Foundation's OpenMDW-1.1 license. NVIDIA claims up to 4x output speed versus similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, and ~670 tokens per second with NVFP4 in pre-release tests, with an Artificial Analysis Intelligence Index of 24 tying gpt-oss-120b. The model is designed for local deployment on hardware like NVIDIA Jetson, GeForce RTX 5090, and DGX Spark, with practical quantized versions fitting on Mac Studio M4 Max 64GB and MacBook Pro M4/M5 Max 48GB, but NVIDIA's speed figures are self-reported and await independent benchmarks.

read3 min views1 publishedAug 12, 2026
Nemotron 3.5 Lightning
Image: Tokenstead (auto-discovered)

MoE workstation**~31.6B total, ~3.6B active per token (MoE)** - hybrid Mamba-Transformer with Multi-Token Prediction (MTP), speculative DSpark/DFlash decoding, and a 1M-token context window.

Released 2026-08-11 as a fully open-weight model under the Linux Foundation’s OpenMDW-1.1 license - weights, data, and recipes on HuggingFace and ModelScope.

Built for fast, long-running agents. NVIDIA positions it as a speed-first engine for always-on agent harnesses (OpenClaw, Hermes Agent, Cline). Claims: up to 4x output speed vs similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, ~670 tok/s with NVFP4 in pre-release tests. Artificial Analysis Intelligence Index 24, tying gpt-oss-120b.

Local-friendly for once. Runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark. BF16 weights ~60GB; the NVFP4 checkpoint is much smaller and the kernels work across Ampere, Hopper, and Blackwell. This is the opposite posture from server-only Nemotron 3 Ultra.

What hardware it actually fits.

Q4_K_M (~32GB weights): the practical quant for most workstations. Fits comfortably on a Mac Studio M4 Max 64GB, MacBook Pro 16-inch M4/M5 Max 48GB, MacBook Pro 14-inch M4/M5 Pro 48GB (with room to spare), and any 64GB+ unified-memory machine. On a 32GB MacBook Pro it is tight - expect memory pressure and swap. - Q8_0 (~58GB weights) / BF16 (~60GB weights): needs 64GB minimum, realistically 80-96GB+ for headroom. Best on Mac Studio M4 Max 96GB, MacBook Pro 16-inch M4/M5 Max 128GB, DGX Spark 128GB, or RTX PRO 6000 Blackwell 96GB. - NVFP4: the intended fast path on NVIDIA Blackwell (RTX 50-series, RTX PRO 6000, DGX Spark GB10). It is not usable on Apple Silicon, AMD, or pre-Blackwell NVIDIA GPUs.

Tokens per second - honest estimates. The ~670 tok/s figure is a pre-release NVIDIA lab number (NVFP4, speculative decode, ideal conditions). Real single-user chat decode is lower. Based on the site’s bandwidth formula and 4K working context, expect roughly:

**MacBook Pro 14-inch M4 Pro 48GB (Q4):** 20-35 t/s -
**MacBook Pro 16-inch M4 Max 48GB / M5 Max 48GB (Q4):** 35-65 t/s -
**Mac Studio M4 Max 64GB (Q4):** 50-75 t/s -
**Mac Studio M4 Max 96GB (Q8/BF16):** 30-45 t/s -
**DGX Spark 128GB (Q4):** 25-40 t/s; (BF16) ~15-25 t/s -
**RTX PRO 6000 Blackwell 96GB (Q4):** 120-180 t/s; (NVFP4) likely 200+ t/s in tuned vLLM -
**RTX 5090 32GB (Q4):** 40-80 t/s, but memory is tight - long context will force offload or crash

These are formula-derived ranges, not measured benchmarks. The model launched August 11, 2026, so community llama.cpp

/ vLLM / Ollama numbers are still pending. Long context, background apps, and thermal throttling will push real numbers toward the lower end of each range.

Honest framing. The speed and agentic numbers are NVIDIA self-reported. The 30B-class size and local deployment claim are the concrete parts; treat the headline throughput figures as claims until independent benchmarks replicate them.

  • 31.6B
  • 1000k
  • other
  • Aug 2026

Scores #

Run it locally #

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

Can you run it? - reference rigs

Rig Q4_K_M Q8_0 FP16
NVIDIA Jetson Orin NX 16GB

no -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

Download options #

Or run it in the cloud #

No per-token API provider pricing tracked for Nemotron 3.5 Lightning yet.

For flagship list prices, see the
[calculator](/calculator).

Inference cost over time #

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nemotron-3-5-lightni…] indexed:0 read:3min 2026-08-12 ·