MoE workstation**~31.6B total, ~3.6B active per token (MoE)** - hybrid Mamba-Transformer with Multi-Token Prediction (MTP), speculative DSpark/DFlash decoding, and a 1M-token context window.
Released 2026-08-11 as a fully open-weight model under the Linux Foundation’s OpenMDW-1.1 license - weights, data, and recipes on HuggingFace and ModelScope.
Built for fast, long-running agents. NVIDIA positions it as a speed-first engine for always-on agent harnesses (OpenClaw, Hermes Agent, Cline). Claims: up to 4x output speed vs similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, ~670 tok/s with NVFP4 in pre-release tests. Artificial Analysis Intelligence Index 24, tying gpt-oss-120b.
Local-friendly for once. Runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark. BF16 weights ~60GB; the NVFP4 checkpoint is much smaller and the kernels work across Ampere, Hopper, and Blackwell. This is the opposite posture from server-only Nemotron 3 Ultra.
What hardware it actually fits.
Q4_K_M (~32GB weights): the practical quant for most workstations. Fits comfortably on a Mac Studio M4 Max 64GB, MacBook Pro 16-inch M4/M5 Max 48GB, MacBook Pro 14-inch M4/M5 Pro 48GB (with room to spare), and any 64GB+ unified-memory machine. On a 32GB MacBook Pro it is tight - expect memory pressure and swap. - Q8_0 (~58GB weights) / BF16 (~60GB weights): needs 64GB minimum, realistically 80-96GB+ for headroom. Best on Mac Studio M4 Max 96GB, MacBook Pro 16-inch M4/M5 Max 128GB, DGX Spark 128GB, or RTX PRO 6000 Blackwell 96GB. - NVFP4: the intended fast path on NVIDIA Blackwell (RTX 50-series, RTX PRO 6000, DGX Spark GB10). It is not usable on Apple Silicon, AMD, or pre-Blackwell NVIDIA GPUs.
Tokens per second - honest estimates. The ~670 tok/s figure is a pre-release NVIDIA lab number (NVFP4, speculative decode, ideal conditions). Real single-user chat decode is lower. Based on the site’s bandwidth formula and 4K working context, expect roughly:
**MacBook Pro 14-inch M4 Pro 48GB (Q4):** 20-35 t/s -
**MacBook Pro 16-inch M4 Max 48GB / M5 Max 48GB (Q4):** 35-65 t/s -
**Mac Studio M4 Max 64GB (Q4):** 50-75 t/s -
**Mac Studio M4 Max 96GB (Q8/BF16):** 30-45 t/s -
**DGX Spark 128GB (Q4):** 25-40 t/s; (BF16) ~15-25 t/s -
**RTX PRO 6000 Blackwell 96GB (Q4):** 120-180 t/s; (NVFP4) likely 200+ t/s in tuned vLLM -
**RTX 5090 32GB (Q4):** 40-80 t/s, but memory is tight - long context will force offload or crash
These are formula-derived ranges, not measured benchmarks. The model launched August 11, 2026, so community llama.cpp
/ vLLM / Ollama numbers are still pending. Long context, background apps, and thermal throttling will push real numbers toward the lower end of each range.
Honest framing. The speed and agentic numbers are NVIDIA self-reported. The 30B-class size and local deployment claim are the concrete parts; treat the headline throughput figures as claims until independent benchmarks replicate them.
- 31.6B
- 1000k
- other
- Aug 2026
Scores #
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
Can you run it? - reference rigs
| Rig | Q4_K_M | Q8_0 | FP16 |
|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB | |||
no -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudno -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Nemotron 3.5 Lightning yet.
For flagship list prices, see the
[calculator](/calculator).
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.