10:03
2026-07-23
promptcube3.com
artificial-intelligence
DGX Spark Inference: 90 tok/s on Large Models
A new inference stack for DGX Spark clusters achieves 55-90 tok/s on large models without speculative decoding, according to internal tests by WoolyAI. The stack enables multi-model agentic workflows โฆ