Hi I am trying to setup a local model machine.
I just got qwen 3 coder 30B A3B and qwen 3.6 27B and 3.8 27B running on my LLM box machine.
I’m new to this and it seems to run okay-ish after applying someone’s advice to quantize and err do something. Lol. I don’t know how it works.
He said it’s new but I could add a second 3090 in my T5810 and run them together over RPC.
I am on a budget so I was thinking of getting a 3080 10GB to put in my server and perhaps it would allow for more context for my qwen’s as it keeps running out.
Also I really don’t know what I’m doing and probably haven’t configured anything properly yet.
I think I get about 70 tokens a second running Qwen 3.8 27B at the moment.
Old mate says he gets 70~ tokens a second too on his 3090.
It just seems misconfigured and slow.
Going to have a second look around the forums for guides on how to configure this while I wait for responses on it. But I thought it was worth asking about dual boxes and RPC config.
My LLM box is 7800X3D, 64GB DDR5, 3090 24GB. Running ubuntu and ollama.
My server is a 14 core intel, 32GB DDR4, 1650 4GB ( at the moment)