cd /news/ai-infrastructure/new-inference-server-for-dgx-spark-l… · home topics ai-infrastructure article
[ARTICLE · art-69309] src=news.ycombinator.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

WoolyAI released a new private multi-agent inference server for DGX Spark clusters, achieving up to 90.83 decode tokens per second on the Nemotron 3 Nano Omni 30B model without quantization or speculative decoding in a 2-node setup. The server enables enterprises to run multi-model agentic workflows cost-effectively by dynamically switching between models, with benchmarks showing 49.30 to 93.31 decode tok/s across DeepSeek V4 Flash, Gemma 4 26B, and Nemotron 3 Nano Omni models. WoolyAI expects to improve performance by 20% with further optimization.

read2 min views1 publishedJul 22, 2026

Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got some really good numbers. Here are the details. The full detailed report is available on our WoolyAI website.

WoolyAI Private Multi-agent Inference Stack for DGX Spark is built to enable groups within enterprises to set up their own private, low-cost inference stacks for business agentic workflow apps. It was built with the forward-looking vision that companies will need their own private, low-cost inference setup and that single-model-based inference is not sufficient for complex enterprise workflow agentic apps. These workflows need multiple models of different specializations (hence sizes) for different steps in the workflows. Having dedicated multi-GPU inference stacks for each model is very cost-prohibitive.

First benchmark: We first ran LlamaBenchy tests on our inference server on a 2 DGX Spark cluster across 3 models with no quantization and no speculative decoding. Using speculative decoding would result in even higher numbers.

MODEL LOAD-PREFILL-SYSTEM-TOK/S DECODE-SYSTEM-TOK/S DECODE-TOK/S-PER-REQUEST DeepSeek V4 Flash C1 1,518.91 21.15 21.15

DeepSeek V4 Flash C4 1,533.15 55.99 14.00

Gemma 4 26B A4B C1 4,579.73 30.22 30.22

Gemma 4 26B A4B C4 4,702.16 63.75 15.94

Nemotron 3 Nano Omni 30B NVFP4 C1 2,607.86 39.42 39.42

Nemotron 3 Nano Omni 30B NVFP4 C4 2,588.87 90.83 22.71

Second Benchmark: One endpoint, three different model-controlled activations. The scheduler batches each burst, coordinates both ranks, and changes the resident model only at a safe boundary

MODEL PREFILL-TOK/S DECODE-TOK/S Model-Activation-Wait DeepSeek V4 Flash 4,154.34 49.30 16s

Gemma 4 26B A4B 18 4,781.44 64.67 6s

Nemotron 3 Nano Omni 30B NVFP4 2,395.59 93.31 2s

We think we can improve these numbers by 20% with more optimization. Please share your feedback. https://woolyai.com/ai-compute-software/dgx-spark-inference-stack/

Comments URL: [https://news.ycombinator.com/item?id=49014048](https://news.ycombinator.com/item?id=49014048)

Points: 1

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @woolyai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/new-inference-server…] indexed:0 read:2min 2026-07-22 ·