00:21
2026-08-22
forum.level1techs.com
large-language-models
Dual Sparks in nvfp4 vs 4x RTX Pro 6000 with native DeepSeek V4 0731 -- Quants and Speed
A homelab user deploying DeepSeek-V4-Flash-0731 on two DGX Spark units achieved 1M token context with an NVFP4 KV cache using a custom vLLM fork, reporting a KV cache size of 1,492,347 tokens and maxi…