Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed Strata v0.1.38 runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a gaming GPU at 94 tokens per second without a server, according to byteiota. The report attributes the feat to the model's MoE architecture and notes that the published benchmarks do not capture everything about the setup. Strata v0.1.38 runs Qwen3.8-Flash-Next — a 125B MoE model — on a gaming GPU at 94 tokens/sec. Here’s how MoE makes it possible, and what the benchmarks don’t tell you. The post Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed appeared first on byteiota .