21:59
2026-07-22
news.ycombinator.com
ai-infrastructure
New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode
WoolyAI released a new private multi-agent inference server for DGX Spark clusters, achieving up to 90.83 decode tokens per second on the Nemotron 3 Nano Omni 30B model without quantization or speculaβ¦