{"slug": "qwen-3-8-27b-on-rtx-5090-at-90-120tps", "title": "Qwen 3.8 27B on RTX 5090 at 90-120tps", "summary": "A developer published a llama-server configuration that runs a Qwen 3.8 27B NVFP4 model with MTP speculative decoding on an RTX 5090, reporting throughput of 90-120 tokens per second. The setup uses a 262,144-token context with q4_0 KV cache quantization, flash attention, and a draft-mtp speculative type capped at four draft tokens.", "body_md": "|  | llama-server \\ | \n|  | --alias qwen3.8-27b-nvfp4-mtp-q8attn \\ | \n|  | -m \"$HOME/.lmstudio/models/utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF/Qwen3.8-27B-NVFP4-MTP-Q8attn.gguf\" \\ | \n|  | --mmproj \"$HOME/.lmstudio/models/utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF/mmproj-Qwen3.8-27B-NVFP4-BF16.gguf\" \\ | \n|  | --spec-type draft-mtp \\ | \n|  | --spec-draft-n-max 4 \\ | \n|  | -ngl 99 \\ | \n|  | -c 262144 \\ | \n|  | -ctk q4_0 \\ | \n|  | -ctv q4_0 \\ | \n|  | -b 2048 \\ | \n|  | -ub 512 \\ | \n|  | --host 127.0.0.1 \\ | \n|  | --port 1234 \\ | \n|  | --jinja \\ | \n|  | --reasoning-effort low \\ | \n|  | --kv-unified \\ | \n|  | -t 12 \\ | \n|  | -np 1 \\ | \n|  | --flash-attn on \\ | \n|  | --no-mmap |", "url": "https://wpnews.pro/news/qwen-3-8-27b-on-rtx-5090-at-90-120tps", "canonical_source": "https://gist.github.com/mikesmullin/c38bba7768671e94b70f65f599155c31", "published_at": "2026-09-22 03:06:05+00:00", "updated_at": "2026-09-22 23:22:52.914999+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "mlops"], "entities": ["Qwen", "RTX 5090", "llama-server", "GGUF", "LM Studio"], "alternates": {"html": "https://wpnews.pro/news/qwen-3-8-27b-on-rtx-5090-at-90-120tps", "markdown": "https://wpnews.pro/news/qwen-3-8-27b-on-rtx-5090-at-90-120tps.md", "text": "https://wpnews.pro/news/qwen-3-8-27b-on-rtx-5090-at-90-120tps.txt", "jsonld": "https://wpnews.pro/news/qwen-3-8-27b-on-rtx-5090-at-90-120tps.jsonld"}}