19:25
2026-08-28
forum.level1techs.com
large-language-models
Qwen 3.8 Flash Next NVFP4 on a single RTX Pro 6000 and system memory
A user benchmarked the Qwen 3.8 Flash Next NVFP4 model on a single RTX Pro 6000 GPU, achieving 81-88 tokens per second in a single session and scaling to 341 tokens per second with six concurrent clieβ¦