# Nemotron 3.5 Lightning

> Source: <https://tokenstead.ai/models/nemotron-3-5-lightning>
> Published: 2026-08-12 12:48:32+00:00

# Nemotron 3.5 Lightning

MoE workstation**~31.6B total, ~3.6B active per token (MoE)** - hybrid Mamba-Transformer with Multi-Token Prediction (MTP), speculative DSpark/DFlash decoding, and a 1M-token context window.

**Released 2026-08-11** as a fully open-weight model under the Linux Foundation’s OpenMDW-1.1 license - weights, data, and recipes on HuggingFace and ModelScope.

**Built for fast, long-running agents.** NVIDIA positions it as a speed-first engine for always-on agent harnesses (OpenClaw, Hermes Agent, Cline). Claims: up to 4x output speed vs similar-sized models, 30% faster task completion on PinchBench than Qwen3.6 35B at similar accuracy, ~670 tok/s with NVFP4 in pre-release tests. Artificial Analysis Intelligence Index 24, tying gpt-oss-120b.

**Local-friendly for once.** Runs on NVIDIA Jetson, GeForce RTX 5090, and DGX Spark. BF16 weights ~60GB; the NVFP4 checkpoint is much smaller and the kernels work across Ampere, Hopper, and Blackwell. This is the opposite posture from server-only Nemotron 3 Ultra.

**What hardware it actually fits.**

-
**Q4_K_M (~32GB weights):** the practical quant for most workstations. Fits comfortably on a Mac Studio M4 Max 64GB, MacBook Pro 16-inch M4/M5 Max 48GB, MacBook Pro 14-inch M4/M5 Pro 48GB (with room to spare), and any 64GB+ unified-memory machine. On a 32GB MacBook Pro it is tight - expect memory pressure and swap. -
**Q8_0 (~58GB weights) / BF16 (~60GB weights):** needs 64GB minimum, realistically 80-96GB+ for headroom. Best on Mac Studio M4 Max 96GB, MacBook Pro 16-inch M4/M5 Max 128GB, DGX Spark 128GB, or RTX PRO 6000 Blackwell 96GB. -
**NVFP4:** the intended fast path on NVIDIA Blackwell (RTX 50-series, RTX PRO 6000, DGX Spark GB10). It is not usable on Apple Silicon, AMD, or pre-Blackwell NVIDIA GPUs.

**Tokens per second - honest estimates.** The ~670 tok/s figure is a pre-release NVIDIA lab number (NVFP4, speculative decode, ideal conditions). Real single-user chat decode is lower. Based on the site’s bandwidth formula and 4K working context, expect roughly:

-
**MacBook Pro 14-inch M4 Pro 48GB (Q4):** 20-35 t/s -
**MacBook Pro 16-inch M4 Max 48GB / M5 Max 48GB (Q4):** 35-65 t/s -
**Mac Studio M4 Max 64GB (Q4):** 50-75 t/s -
**Mac Studio M4 Max 96GB (Q8/BF16):** 30-45 t/s -
**DGX Spark 128GB (Q4):** 25-40 t/s; (BF16) ~15-25 t/s -
**RTX PRO 6000 Blackwell 96GB (Q4):** 120-180 t/s; (NVFP4) likely 200+ t/s in tuned vLLM -
**RTX 5090 32GB (Q4):** 40-80 t/s, but memory is tight - long context will force offload or crash

These are formula-derived ranges, not measured benchmarks. The model launched August 11, 2026, so community `llama.cpp`

/ vLLM / Ollama numbers are still pending. Long context, background apps, and thermal throttling will push real numbers toward the lower end of each range.

**Honest framing.** The speed and agentic numbers are NVIDIA self-reported. The 30B-class size and local deployment claim are the concrete parts; treat the headline throughput figures as claims until independent benchmarks replicate them.

- 31.6B
- 1000k
- other
- Aug 2026

## Scores

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | Q4_K_M | Q8_0 | FP16 |
|---|---|---|---|
| NVIDIA Jetson Orin NX 16GB |
|

[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

No per-token API provider pricing tracked for Nemotron 3.5 Lightning yet.
For flagship list prices, see the
[calculator](/calculator).

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
