Strata v0.1.38 runs Qwen3.8-Flash-Next — a 125B MoE model — on a gaming GPU at 94 tokens/sec. Here’s how MoE makes it possible, and what the benchmarks don’t tell you.
The post Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed appeared first on byteiota .