cd /news/large-language-models/strata-runs-a-125b-ai-model-on-your-… · home › topics › large-language-models › article
[ARTICLE · art-145188] src=byteiota.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed

Strata v0.1.38 runs Qwen3.8-Flash-Next, a 125B mixture-of-experts model, on a gaming GPU at 94 tokens per second without a server, according to byteiota. The report attributes the feat to the model's MoE architecture and notes that the published benchmarks do not capture everything about the setup.

read1 min views1 publishedOct 5, 2026

Strata v0.1.38 runs Qwen3.8-Flash-Next — a 125B MoE model — on a gaming GPU at 94 tokens/sec. Here’s how MoE makes it possible, and what the benchmarks don’t tell you.

The post Strata Runs a 125B AI Model on Your Gaming PC — No Server Needed appeared first on byteiota .

── more in #large-language-models 4 stories · sorted by recency
── more on @strata 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/strata-runs-a-125b-a…] indexed:0 read:1min 2026-10-05 · —