{"slug": "new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode", "title": "New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode", "summary": "WoolyAI released a new private multi-agent inference server for DGX Spark clusters, achieving up to 90.83 decode tokens per second on the Nemotron 3 Nano Omni 30B model without quantization or speculative decoding in a 2-node setup. The server enables enterprises to run multi-model agentic workflows cost-effectively by dynamically switching between models, with benchmarks showing 49.30 to 93.31 decode tok/s across DeepSeek V4 Flash, Gemma 4 26B, and Nemotron 3 Nano Omni models. WoolyAI expects to improve performance by 20% with further optimization.", "body_md": "Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.\n\nWoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.\n\nFirst benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.\n\nMODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST\n\nDeepSeek V4 Flash C1 1,518.91 21.15 21.15\n\nDeepSeek V4 Flash C4 1,533.15 55.99 14.00\n\nGemma 4 26B A4B C1 4,579.73 30.22 30.22\n\nGemma 4 26B A4B C4 4,702.16 63.75 15.94\n\nNemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42\n\nNemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71\n\nSecond Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary\n\nMODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait\n\nDeepSeek V4 Flash 4,154.34 49.30 16s\n\nGemma 4 26B A4B 18 4,781.44 64.67 6s\n\nNemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s\n\nWe think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/\n\nComments URL: [https://news.ycombinator.com/item?id=49014048](https://news.ycombinator.com/item?id=49014048)\n\nPoints: 1\n\n# Comments: 0", "url": "https://wpnews.pro/news/new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode", "canonical_source": "https://news.ycombinator.com/item?id=49014048", "published_at": "2026-07-22 21:59:39+00:00", "updated_at": "2026-07-22 22:22:33.603336+00:00", "lang": "en", "topics": ["ai-infrastructure", "ai-products", "ai-tools", "machine-learning"], "entities": ["WoolyAI", "DGX Spark", "DeepSeek V4 Flash", "Gemma 4 26B", "Nemotron 3 Nano Omni"], "alternates": {"html": "https://wpnews.pro/news/new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode", "markdown": "https://wpnews.pro/news/new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode.md", "text": "https://wpnews.pro/news/new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode.txt", "jsonld": "https://wpnews.pro/news/new-inference-server-for-dgx-spark-large-model-c4-55-90-tok-s-no-spec-decode.jsonld"}}