Strata: Qwen3.8-Flash-Next (125B Moe) on a 8GB+ Nvidia GPU Strata, an open-source tool from developer Niko1221, runs the 125-billion-parameter Qwen3.8-Flash-Next model on a single NVIDIA GPU with 12-24 GB of VRAM and 64 GB of RAM, generating answers at 60-95 tokens per second. Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM, the Q2_0 quantization writes at 90 tokens/s in short chat and 67 tokens/s at 128K context, while IQ3_S reaches 52 and 41 tokens/s respectively. The release also packages ISTA-DASLab's Coder variant, which keeps 91% of the full model's SWE-bench Verified score and 99% of LiveCodeBench, and UkisAI's Swift 1.5 fine-tune. Run a 125-billion-parameter AI model on a normal gaming PC one NVIDIA card 12-24 GB + 64 GB of RAM · Windows or Linux · one click to install