{"slug": "sanity-check-my-qwen3-8-5090-results-new-to-local-ai", "title": "Sanity check my Qwen3.8 5090 Results; new to Local AI", "summary": "A user running Qwen3.8-27B on an RTX 5090 with the ninfer runtime reported a 1.72× speedup over LM Studio/llama.cpp, achieving 124.08 tok/s decode and 8,143.82 tok/s prompt eval with mixed NVFP4/FP8 quantization, versus 72.03 tok/s and 2,430.18 tok/s with Q6_K. The user noted that token speeds vary by output type and suggested comparing identical prompts for accurate benchmarking.", "body_md": "Here’s a datapoint from my 5090; I’ve been running qwen3.8 with [GitHub - Neroued/ninfer: High-performance single-GPU inference for selected model checkpoints and GPUs. · GitHub](https://github.com/Neroued/ninfer) . Similar to you some of this local model stuff has outpaced me since I last checked in, don’t completely grasp the impact of MTP and KVcache quant yet. But this ninfer seemed to have some optimizations for single gpu 5090 arch, but it may not work with the finetunes you need. I also pulled one of the uncensored models with lmstudio ran with llama-benchy as a comparison.\n\nspecs\n\n9950x3d\n\n5090\n\n96GB 6400mhz\n\nlinux\n\n| Model / Runtime | Quant | MTP Max | Decode | Prompt Eval | Approx. VRAM | Relative Speed |\n|---|---|---|---|---|---|---|\n| JonathanColetti/Qwen3.8-27B-Uncensored-GGUF via LM Studio/llama.cpp | Q6_K | 2 | 72.03 ± 5.66 tok/s | 2,430.18 ± 24.45 tok/s | ~28.5 GB total GPU use | 1.00× |\n| neroued/Qwen3.8-27B-nvfp4-NInfer via NInfer | mixed NVFP4/FP8 | 3 | 124.08 ± 10.56 tok/s | 8,143.82 ± 52.68 tok/s | ~28.5 GB total GPU use | 1.72× |\n\nLooks like these models tok/s can vary pretty heavily based on the type of output being generated. So, hard to say how useful this will be for you without comparing identical prompts. I’m interested in your usecase though, uncensored model + red teaming on 5090. Let me know if you find some optimal config with vLLM", "url": "https://wpnews.pro/news/sanity-check-my-qwen3-8-5090-results-new-to-local-ai", "canonical_source": "https://forum.level1techs.com/t/sanity-check-my-qwen3-8-5090-results-new-to-local-ai/254299#post_4", "published_at": "2026-08-24 07:58:06+00:00", "updated_at": "2026-08-24 10:44:04.384888+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["Qwen3.8-27B", "RTX 5090", "ninfer", "LM Studio", "llama.cpp", "JonathanColetti", "Neroued"], "alternates": {"html": "https://wpnews.pro/news/sanity-check-my-qwen3-8-5090-results-new-to-local-ai", "markdown": "https://wpnews.pro/news/sanity-check-my-qwen3-8-5090-results-new-to-local-ai.md", "text": "https://wpnews.pro/news/sanity-check-my-qwen3-8-5090-results-new-to-local-ai.txt", "jsonld": "https://wpnews.pro/news/sanity-check-my-qwen3-8-5090-results-new-to-local-ai.jsonld"}}