I moved my BOSGAME M5 from Windows 11 to Ubuntu, and it got a lot faster at local AI. A long Qwen3.8 answer now takes 25.3 seconds instead of 37.9, and the first text arrives 11 seconds sooner. The ASUS GX10 is still quicker, but the gap is much smaller now.
In my original review, the M5 lost the long-prompt Qwen3.8 test to my ASUS GX10, mostly because of the long wait before the first token. I wanted to see how much faster it gets on Linux, so I installed Ubuntu and ran the same test again. The complete answer now takes 25.27 seconds instead of 37.94 seconds. That is 33.38% less time on the same machine.
This is the M5 from my original review against the ASUS GX10: a Ryzen AI Max+ 395 with Radeon 8060S graphics and 128GB of shared memory. The Windows and GX10 numbers in this article come from that review’s AI test page.
I am getting the M5 ready to take over Cybi and Buster, our public data assistants, and Cybenetics AI Check. The assistants need fast answers. AI Check can take longer, as long as it gets the report right and leaves the assistants room to work. Here, I cover the AI speed, the temperatures, and the memory on Ubuntu, and what I still have to test before the move.
Related BOSGAME M5 articles:
- [BOSGAME M5 AI PC Review: Strix Halo vs My 4TB ASUS GX10](https://hwbusters.com/systems/bosgame-m5-ai-pc-review-strix-halo-vs-asus-gx10/)
- [Bosgame M5 Fan Control: Broken as Shipped, So I Fixed It With a Free Plugin](https://hwbusters.com/systems/bosgame-m5-fan-control-broken-as-shipped-so-i-fixed-it-with-a-free-plugin/)
Test Setup #
I installed Ubuntu 26.04.1 with the Xfce desktop and kept the 96GB graphics allocation I used on Windows, so the large models fit in GPU memory.
| Configuration | BOSGAME M5 on Ubuntu | | Operating system | Ubuntu 26.04.1, kernel 7.0.0-38, Xfce desktop | | Graphics driver | Mesa/RADV 26.0.8 | | Inference engine | llama.cpp b11475, Vulkan and HIP/ROCm builds | | ROCm | Separate ROCm 10.1 gfx1151 userspace; binary built against ROCm 10.0 | | Graphics memory | 96GB, set in the BIOS | | Public-chat model | Qwen3 30B-A3B Q8 | | Report model | Qwen3.8 27.3B Q8 | | Context | 8,192 tokens in every test | | Fans | Both main fans fixed at 80% for the benchmarks |
I ran every test on both of llama.cpp’s GPU backends: Vulkan, through the Mesa driver, and HIP/ROCm, through a separate ROCm 10.1 install for the gfx1151 GPU.
Two models do the work. Qwen3 30B-A3B Q8, a mixture-of-experts model, answers the public chat. Qwen3.8 27.3B Q8 checks the reports, and it is also my candidate for the SQL work. They run different workloads, so I compare each model only with itself.
For the public-chat tests, I used exact input and output token counts, one warm-up, and five measured requests per case. The Qwen3.8 test repeats my original review’s method, with three natural answers per scenario.