{"slug": "strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed", "title": "Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed", "summary": "Strata v0.1.38 runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a gaming GPU at 94 tokens per second without a server, according to byteiota. The report attributes the feat to the model's MoE architecture and notes that the published benchmarks do not capture everything about the setup.", "body_md": "Strata v0.1.38 runs Qwen3.8-Flash-Next — a 125B MoE model — on a gaming GPU at 94 tokens/sec. Here’s how MoE makes it possible, and what the benchmarks don’t tell you.\n\nThe post \nStrata Runs a 125B AI Model on Your Gaming PC — No Server Needed\n appeared first on \nbyteiota\n.", "url": "https://wpnews.pro/news/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed", "canonical_source": "https://byteiota.com/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed/", "published_at": "2026-10-05 05:08:41+00:00", "updated_at": "2026-10-05 05:11:59.576321+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "machine-learning"], "entities": ["Strata", "Strata v0.1.38", "Qwen3.8-Flash-Next", "byteiota"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed", "markdown": "https://wpnews.pro/news/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed.md", "text": "https://wpnews.pro/news/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed.txt", "jsonld": "https://wpnews.pro/news/strata-runs-a-125b-ai-model-on-your-gaming-pc-no-server-needed.jsonld"}}