{"slug": "sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai", "title": "Sanitfy check my Qwen3.8 5090 Results; new to Local AI", "summary": "A user running Qwen3.8-27B on an NVIDIA RTX 5090 32 GB with LM Studio reports decode speeds ranging from 12.38 to 47.4 tok/s depending on quantization and multi-token prediction settings, with the Q4_K_M quant achieving the fastest speed at ~47.4 tok/s and ~22.9 GB VRAM usage. The user, who works in AI adoption, governance, and cybersecurity, is building a 20-40b MCP server for red/blue teaming with Kali and seeks feedback on whether their setup and results are optimal.", "body_md": "Hey Everyone! First post, kinda shy, be gentle\n\nSo - I don’t have a TON of experience in Local AI - I do work in the AI Adoption / Governance / Cybersecurity Space, and am lucky enough to have active subscriptions across every major frontier model - however, I haven’t messed with Local AI in about 8 months. But - I’ve been hearing a ton about Qwen3.8-27b, so I loaded it up on my rig at the house in LM Studio, and ran a few tests with a harness doing some pentesting on a dev vm, and was pretty happy with the initial results, though it did get stuck in a few loops which ate through my context. So - I’m now in the process of building a 20-40b focused MCP server specifically for red teaming / blue teaming with Kali.\n\nBefore I get too deep in that process (I’m around 15ish hours in so far) - I want to get some opinions on if my initial testing is “Stupid” or “The old way”.\n\nThis is my current setup:\n\n| HW | Spec |\n|---|---|\n| CPU | Intel Core i9-13900K |\n| GPU | NVIDIA RTX 5090 32 GB |\n| RAM | 64 GB 7200 MT/s DDR5 |\n| OS | Windows |\n\nI’m using LM Studio right now - which I’m sure isn’t great - but this is the only thing I’ve used for tasks like these - open to suggestions if this is the wrong setup.\n\nBelow is the performance I have been getting - this is another area I want a sanity check on. Before that, a bit of context for my decisions - this test was scoped specifically for a phased, agentic red team assessment on a single device. Below are my results:\n\n| Model / GGUF | Quant | MTP Max | Decode | Prompt Eval | Approx. VRAM | Relative Speed |\n|---|---|---|---|---|---|---|\n`chimingw/Qwen3.8-27B-Uncensored-OrcaRouter-GGUF` |\nQ6_K | 2 | 12.38 tok/s |\n1,016.55 tok/s |\n~28 GB | 1.00× |\n`JonathanColetti/Qwen3.8-27B-Uncensored-GGUF` |\nQ6_K | 1 | ~22.3 tok/s |\n— | ~27.5–28.5 GB | 1.80× |\n`JonathanColetti/Qwen3.8-27B-Uncensored-GGUF` |\nQ6_K | 2 | 31.60 tok/s |\n~1,844.91 tok/s |\n~27.5–28.5 GB | 2.55× |\n`JonathanColetti/Qwen3.8-27B-Uncensored-GGUF` |\nQ5_K_M | 2 | ~44.7 tok/s |\n~2,045.19 tok/s |\n~26.3–26.5 GB |\n3.61× |\n`JonathanColetti/Qwen3.8-27B-Uncensored-GGUF` |\nQ4_K_M | 2 | ~47.4 tok/s |\n~2,267.78 tok/s |\n~22.9 GB |\n3.83× |\n\nSo - open ended question; give me your feedback. Should I switch away from LM Studio? Are my speeds decent? Am I leaving some performance on the table?", "url": "https://wpnews.pro/news/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai", "canonical_source": "https://forum.level1techs.com/t/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai/254299#post_1", "published_at": "2026-08-23 19:09:11+00:00", "updated_at": "2026-08-23 19:13:26.841431+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["Qwen3.8-27B", "LM Studio", "NVIDIA RTX 5090", "Intel Core i9-13900K", "Kali", "MCP"], "alternates": {"html": "https://wpnews.pro/news/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai", "markdown": "https://wpnews.pro/news/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai.md", "text": "https://wpnews.pro/news/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai.txt", "jsonld": "https://wpnews.pro/news/sanitfy-check-my-qwen3-8-5090-results-new-to-local-ai.jsonld"}}