05:08
2026-10-05
byteiota.com
large-language-models
Strata Runs a 125B AI Model on Your Gaming PC โ No Server Needed
Strata v0.1.38 runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a gaming GPU at 94 tokens per second without a server, according to byteiota. The report attributes the feat to the model's โฆ