15:07
2026-09-28
github.com
large-language-models
Strata: Qwen3.8-Flash-Next (125B Moe) on a 8GB+ Nvidia GPU
Strata, an open-source tool from developer Niko1221, runs the 125-billion-parameter Qwen3.8-Flash-Next model on a single NVIDIA GPU with 12-24 GB of VRAM and 64 GB of RAM, generating answers at 60-95 …