22:04
2026-08-14
gist.github.com
large-language-models
Qwen3.8-27B NVFP4-MTP benchmark on RTX 5090 (tier comparison, 64K long-context, spec-draft-n-max sweep) + mtp-bench.py tool
A developer benchmarked the Qwen3.8-27B model with NVFP4 quantization and multi-token prediction (MTP) on an RTX 5090, comparing tier configurations and sweeping spec-draft-n values at 64K context. Thβ¦