21:45
2026-08-26
gist.github.com
large-language-models
Qwen3.8-Flash-Next UD-Q4_K_XL on one DGX Spark: 128K ctx + q8 KV (fix for qwen4exp self_k_rot assert)
A developer tested the Qwen3.8-Flash-Next model on an NVIDIA DGX Spark (GB10, 128 GB unified memory) using llama.cpp with quantized KV cache, achieving 128K context and 414 tok/s prefill. They identifβ¦