18:59
2026-08-20
alephinitesimal.com
artificial-intelligence
Benchmarking Speculative Decoding on DGX Spark: 11.5 β 29.5 tok/s
NVIDIA's DGX Spark, powered by the GB10 superchip with 128 GB of unified LPDDR5X memory, achieves only 11.5 tokens per second running Qwen3.8-27B at Q4_K_XL quantization, limited by 273 GB/s memory baβ¦