{"slug": "dgx-spark-inference-90-tok-s-on-large-models", "title": "DGX Spark Inference: 90 tok/s on Large Models", "summary": "A new inference stack for DGX Spark clusters achieves 55-90 tok/s on large models without speculative decoding, according to internal tests by WoolyAI. The stack enables multi-model agentic workflows across a 2 DGX Spark cluster, with LlamaBench results showing 55.99 tok/s for DeepSeek V4 Flash C4, 63.75 tok/s for Gemma 4 26B A4B C4, and 90.83 tok/s for Nemotron 3 Nano Omni 30B NVFP4 C4. A 20% performance improvement is expected after the next optimization round.", "body_md": "# DGX Spark Inference: 90 tok/s on Large Models\n\nHitting 55-90 tok/s on large models without even touching speculative decoding is a massive win for our current deployment. We've been testing a new inference stack specifically for DGX Spark clusters to handle multi-model agentic workflows, and the LlamaBench numbers are actually holding up under simulated traffic.\n\nWe also tested a scenario with one endpoint controlling three different model activations. The scheduler batches bursts and coordinates ranks, only swapping the resident model at safe boundaries. The results were:\n\nThis is a solid deep dive into how we're scaling our AI workflow internally. We're expecting another 20% bump in performance once we finish the next round of optimization.\n\nThe real-world problem we're solving at work is that enterprise agentic apps usually need a mix of specialized models of different sizes. Buying dedicated multi-GPU stacks for every single model in a workflow is way too expensive. This setup lets us run a private, lower-cost stack that handles multiple models across a 2 DGX Spark cluster.\n\nHere is the raw performance data from the LlamaBenchy tests (no quantization, no spec decode):\n\n**DeepSeek V4 Flash C4:** 55.99 decode-system-tok/s**Gemma 4 26B A4B C4:** 63.75 decode-system-tok/s**Nemotron 3 Nano Omni 30B NVFP4 C4:** 90.83 decode-system-tok/s\n\nWe also tested a scenario with one endpoint controlling three different model activations. The scheduler batches bursts and coordinates ranks, only swapping the resident model at safe boundaries. The results were:\n\n**DeepSeek V4 Flash:** 49.30 decode-tok/s (16s activation wait)**Gemma 4 26B A4B:** 64.67 decode-tok/s (6s activation wait)**Nemotron 3 Nano Omni 30B NVFP4:** 93.31 decode-tok/s (2s activation wait)\n\nThis is a solid deep dive into how we're scaling our AI workflow internally. We're expecting another 20% bump in performance once we finish the next round of optimization.\n\nFor those wanting the full technical breakdown, the report is here:`https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/`\n\n[Next Framework Desktop Build: Ryzen AI Max+ Pro 495 & 192GB RAM →](/en/threads/2194/)\n\n## All Replies （4）\n\nN\n\nSaw similar jumps after switching my stack. Memory bandwidth is usually the hidden bottleneck here.\n\n0\n\nC\n\nI noticed a decent boost using FP8 quantization too, might be worth a quick test.\n\n0\n\nS\n\n[@CyberSmith](/en/users/CyberSmith/)Does FP8 hit the accuracy hard or is it pretty seamless? I've been scared to try it.\n\n0\n\nS\n\nDid you see any stability issues or spikes in latency during longer context windows?\n\n0", "url": "https://wpnews.pro/news/dgx-spark-inference-90-tok-s-on-large-models", "canonical_source": "https://promptcube3.com/en/threads/2213/", "published_at": "2026-07-23 10:03:34+00:00", "updated_at": "2026-07-23 18:41:01.630458+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-infrastructure", "ai-products", "ai-tools"], "entities": ["WoolyAI", "DGX Spark", "DeepSeek V4 Flash C4", "Gemma 4 26B A4B C4", "Nemotron 3 Nano Omni 30B NVFP4 C4", "LlamaBench"], "alternates": {"html": "https://wpnews.pro/news/dgx-spark-inference-90-tok-s-on-large-models", "markdown": "https://wpnews.pro/news/dgx-spark-inference-90-tok-s-on-large-models.md", "text": "https://wpnews.pro/news/dgx-spark-inference-90-tok-s-on-large-models.txt", "jsonld": "https://wpnews.pro/news/dgx-spark-inference-90-tok-s-on-large-models.jsonld"}}