{"slug": "running-qwen-3-6-2-8-on-3090-3080-over-rpc", "title": "Running qwen 3.6 / 2.8 on 3090+3080 over RPC?", "summary": "A user reports running Qwen 3 Coder 30B A3B, Qwen 3.6 27B, and Qwen 3.8 27B on a local machine with a 7800X3D, 64GB DDR5, and an RTX 3090 24GB, achieving about 70 tokens per second on Qwen 3.8 27B, and is considering adding an RTX 3080 10GB to a server for more context and dual-box RPC setup, though they admit to being new and possibly misconfigured.", "body_md": "Hi I am trying to setup a local model machine.\n\nI just got qwen 3 coder 30B A3B and qwen 3.6 27B and 3.8 27B running on my LLM box machine.\n\nI’m new to this and it seems to run okay-ish after applying someone’s advice to quantize and err do something. Lol. I don’t know how it works.\n\nHe said it’s new but I could add a second 3090 in my T5810 and run them together over RPC.\n\nI am on a budget so I was thinking of getting a 3080 10GB to put in my server and perhaps it would allow for more context for my qwen’s as it keeps running out.\n\nAlso I really don’t know what I’m doing and probably haven’t configured anything properly yet.\n\nI think I get about 70 tokens a second running Qwen 3.8 27B at the moment.\n\nOld mate says he gets 70~ tokens a second too on his 3090.\n\nIt just seems misconfigured and slow.\n\nGoing to have a second look around the forums for guides on how to configure this while I wait for responses on it. But I thought it was worth asking about dual boxes and RPC config.\n\nMy LLM box is 7800X3D, 64GB DDR5, 3090 24GB. Running ubuntu and ollama.\n\nMy server is a 14 core intel, 32GB DDR4, 1650 4GB ( at the moment)", "url": "https://wpnews.pro/news/running-qwen-3-6-2-8-on-3090-3080-over-rpc", "canonical_source": "https://forum.level1techs.com/t/running-qwen-3-6-2-8-on-3090-3080-over-rpc/254271#post_1", "published_at": "2026-08-23 03:27:30+00:00", "updated_at": "2026-08-23 03:42:52.647956+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure"], "entities": ["Qwen", "Qwen 3 Coder 30B A3B", "Qwen 3.6 27B", "Qwen 3.8 27B", "RTX 3090", "RTX 3080", "Ollama", "Ubuntu"], "alternates": {"html": "https://wpnews.pro/news/running-qwen-3-6-2-8-on-3090-3080-over-rpc", "markdown": "https://wpnews.pro/news/running-qwen-3-6-2-8-on-3090-3080-over-rpc.md", "text": "https://wpnews.pro/news/running-qwen-3-6-2-8-on-3090-3080-over-rpc.txt", "jsonld": "https://wpnews.pro/news/running-qwen-3-6-2-8-on-3090-3080-over-rpc.jsonld"}}