{"slug": "strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu", "title": "Strata: Qwen3.8-Flash-Next (125B Moe) on a 8GB+ Nvidia GPU", "summary": "Strata, an open-source tool from developer Niko1221, runs the 125-billion-parameter Qwen3.8-Flash-Next model on a single NVIDIA GPU with 12-24 GB of VRAM and 64 GB of RAM, generating answers at 60-95 tokens per second. Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM, the Q2_0 quantization writes at 90 tokens/s in short chat and 67 tokens/s at 128K context, while IQ3_S reaches 52 and 41 tokens/s respectively. The release also packages ISTA-DASLab's Coder variant, which keeps 91% of the full model's SWE-bench Verified score and 99% of LiveCodeBench, and UkisAI's Swift 1.5 fine-tune.", "body_md": "**Run a 125-billion-parameter AI model on a normal gaming PC**\n\none NVIDIA card (12-24 GB) + 64 GB of RAM · Windows or Linux · one click to install\n\n<sub>A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) ·\n[full video (49 s)](https://github.com/Niko1221/Strata/releases/download/v0.1.10/Pagoda.mp4)</sub>\n\nStrata runs **[Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next)** - a large, smart AI model that\nnormally needs a server - on your own PC. It writes its answers at **60-95 tokens per second** (a token is about ¾\nof a word): faster than you can read.\n\n- **Free and open source.**\n\n**Jump to:** [How fast?](#how-fast-is-it) · [Which model?](#which-model-should-i-pick) · [Install](#install) ·\n[Using it](#using-it) · [Problems?](#something-went-wrong) · [How it works](#how-does-it-work) ·\n[All the details](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md)\n\nMeasured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:\n\n| Size | Writes answers (short chat) | Writes answers (128K context) | Reads your prompt | \n|---|---|---|---|\n| **Q2_0** | 90 tokens/s | 67 tokens/s | 1,310 tokens/s | \n| **IQ2_XS** | 74 tokens/s | 60 tokens/s | 1,240 tokens/s | \n| **IQ3_XXS** | 62 tokens/s | 46 tokens/s | 1,110 tokens/s | \n| **IQ3_S** | 52 tokens/s | 41 tokens/s | 1,070 tokens/s | \n| **Coder** (IQ1_M) | 51 tokens/s | 44 tokens/s | 1,300 tokens/s | \n\n- **Writes answers** = how fast the reply appears (tokens per second).\n- **Reads your prompt** = how fast it takes in what you send (long documents, code, chat history), measured on a\n32K-token prompt; a 4K prompt reads at 740-1,000 tokens/s. A 32K prompt takes about 25 seconds with Q2_0.\n\nA card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly\n100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the\n[details](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#speed-measured).\n\n**The size** (the same model, compressed more or less):\n\n| Model | RAM+VRAM Requirements | Speed | Quality | \n|---|---|---|---|\n| **Q2_0** | 37.6 GB | fastest | good | \n| **IQ2_XS** | 39.2 GB | fast | better ( **recommended** ) | \n| **IQ3_XXS** | 47.0 GB | slower | great | \n| **IQ3_S** | 54.8 GB | slowest | best: matches the full model on the published tests (original model only) | \n\n**Will it fit?** Shard 1 is the part of the model that gets loaded when it starts: its experts go into your **RAM**,\nthe rest onto your graphics card (the second shard, a 29 GB lookup table, stays on the SSD). So it fits when your\n**RAM is at least shard 1 + about 10 GB** for Windows and your other programs. With 64 GB of RAM every size fits\n(IQ3_S with little else open); with 48 GB, Q2_0 and IQ2_XS. A bigger graphics card makes it faster, but it doesn't\nlower the RAM needed.\n\n**The version:**\n\n- **Qwen3.8-Flash-Next** - the original.\n- **[Coder](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF)** - ISTA-DASLab's coding\nversion: half of the experts removed, keeping the ones that code, tool use and images need (91% of the full model's\nSWE-bench Verified score, 99% of LiveCodeBench, by its authors). One size (IQ1_M: its experts stored like IQ3_S):\nshard 1 is**29.6 GB** , so it fits a PC with**32 GB of RAM** , runs 262K context on 64 GB, and reads long prompts\nthe fastest of all. Weaker outside coding.\n- **[Swift 1.5](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF)** - a fine-tune by UkisAI\nthat thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per\ntoken, and about the same RAM as the same size of the original (no IQ3_S). Its own license applies (see its page).\n\nNot sure? Take **IQ2_XS** - or the **Coder** if you mainly write code, or have 32-48 GB of RAM. You can add another\none later with `SETUP.bat` (the same as `START-HERE.bat --setup`; on Linux `./setup.sh --setup`).\n\n**You need:** an NVIDIA RTX 30, 40 or 50 card with 12 GB of VRAM or more, enough RAM for the size you pick (above),\n~80 GB of free disk space (an SSD makes the first start much faster), and Windows 10/11 or Linux. The only thing you\ninstall yourself is a current **NVIDIA driver** ([nvidia.com/drivers](https://www.nvidia.com/drivers) or the NVIDIA\nApp). Everything else - Python, the engine, the model - is set up for you.\n\n**Windows**\n\n1. [Download this project](https://github.com/Niko1221/Strata/archive/refs/heads/main.zip) and unzip it (or`git clone` it).\n2. Double-click **`START-HERE.bat`** .\n3. Answer a few questions - or just press Enter each time for the recommended choice:\n  - **Which model and size?** The original or Swift 1.5, and Q2_0, IQ2_XS, IQ3_XXS or IQ3_S - see[above](#which-model-should-i-pick)\n  - **How much context?** How much text it can keep in mind at once (it suggests one for your card)\n  - **Images?** Whether it should also read pictures\n  - **Experimental speed projection?** Off unless you say yes -[read what it does](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#experimental-speed-projection-experimental-off-by-default) first\n\nThen it downloads everything (the model is ~70 GB, so the first time takes a while - you can stop and it picks up\nwhere it left off) and **starts the model**. Your browser opens the Strata app at `http://127.0.0.1:8080`.\n\n**While the model starts, your PC can be slow or stop responding for 1-3 minutes** (longest the first time): Strata\nloads 35-55 GB into your RAM and locks part of it for the graphics card. That's normal - wait, and don't close the\nwindow. The window tells you what it is doing.\n\n**Next time**, just double-click `START-HERE.bat` again: it starts right away, nothing is downloaded twice. Close its\nwindow to stop the model.\n\n**Updating:** download the new version and unzip it anywhere (or `git pull`), then run `START-HERE.bat` in it. The\nmodel files are kept in a `Strata-data` folder next to your Strata folder, so a new copy finds them and sets itself up\nthe same way - nothing big is downloaded again.\n\n**Linux:** run `./setup.sh` - same questions, same result.\n\n<sub>The Strata app's **Monitor** (left) while a coding agent writes the pagoda garden from the video (right)</sub>\n\n- **In the browser:**`http://127.0.0.1:8080` - the Strata app (it opens by itself when the model starts):** Chat** , a\nlive**Monitor** of the model and your GPU/CPU/RAM, and**About** with the settings and addresses.\n- **Chat in the terminal:**`.venv\\Scripts\\python chat.py`\n- **Your apps and coding agents:** add it as an \"OpenAI-compatible\" provider with base URL**`http://127.0.0.1:8080/v1`** , any API key and any model name. Apps that use Anthropic's API:`http://127.0.0.1:8080/v1/messages` .\n- **Thinking:** the model thinks before it answers. Choose**off, low, medium or high** - in the chat page menu, with`/think low` in`chat.py` , or with your app's \"reasoning effort\" setting. Off is fastest; high is best for hard questions.\n- **Pictures:** in the chat page click**Picture** ; in`chat.py` type`/image <path>` ; in apps just attach them.\n- **From your phone or another PC:**`START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>` , then open the\naddress the server window prints; see the[details](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#using-it) .\n- **Experimental speed projection (off by default):** an experimental control vector that setup can turn on; it\nchanges how the model answers - read[what it does](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#experimental-speed-projection-experimental-off-by-default) first.\n\n**Good to know:** it answers one request at a time. The first message of a chat is read in full (about 1 minute per\n30,000 tokens); after that it keeps the conversation and reads only what is new, so follow-ups start in seconds.\n\n**My PC froze, or got very slow, the first time Strata started.**\nThat's normal while it starts, most of all the first time. Strata loads 35-55 GB into your RAM, locks part of it for\nthe graphics card, and works out how much of the model fits on your GPU. The mouse can freeze for a few minutes. **Wait, and don't close the\nwindow.** The next starts are much faster. Still frozen after 10 minutes? Restart the PC, close other programs\n(browsers use a lot of RAM) and try again. If it keeps happening, pick a smaller size (Q2_0 or IQ2_XS).\n\n**It stopped while downloading or installing.**\nRun `START-HERE.bat` again. It continues where it stopped.\n\n**It says the NVIDIA driver is too old.**\nUpdate it (NVIDIA App or [nvidia.com/drivers](https://www.nvidia.com/drivers)), restart the PC, and run\n`START-HERE.bat` again.\n\n**It says port 8080 is already in use.**\nStrata is already running. Look for its window.\n\n**It's very slow and the disk light keeps blinking.**\nYour PC is out of free RAM. Close other programs, or pick a smaller size (Q2_0 or IQ2_XS).\n\n**An answer stopped with \"the engine stopped unexpectedly\".**\nUsually not enough RAM (on Linux the system then stops the engine). Just send your message again: Strata starts the\nengine by itself. If it keeps happening, close other programs or pick a smaller size.\n\n**It says the prompt exceeds the context.**\nThe conversation is longer than the context you chose. Start a new chat, or run `SETUP.bat` and pick more\ncontext.\n\n**Still stuck?** Look in the [full troubleshooting table](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#troubleshooting), or open an issue and\nattach `strata-<model>.log` from the Strata folder.\n\nModels like this one normally run on servers with hundreds of gigabytes of graphics memory. Your graphics card has\n12-24 GB. Strata makes it fit by **sharing the work across your whole PC** - the same idea as a kitchen, where the\nthings you use all the time stay on the counter and the rest waits in the pantry.\n\n- **The model is a team of 24,576 small specialists (\"experts\"),** and each word it writes needs only 10 of them.\nSo it doesn't have to have all of them on the graphics card at once.\n- **Your graphics card** does the part of the work needed for every word, and keeps the few thousand experts that\nare asked most often. It keeps learning which ones those are while you use it.\n- **Your RAM** holds every expert. When a word needs one the card doesn't have,**your processor** works on it -\nat the same time as the graphics card, so neither waits for the other.\n- **Your SSD** holds a big lookup table; the model only reads a few small rows of it per word.\n\n- **Guess, then check.** A small, fast helper built into the model guesses the next few words, and the big model\nchecks all the guesses in one go. It keeps the ones it agrees with and writes the next word itself - so one step\noften produces several words. The helper only guesses - the big model decides every word - so you get the same\nquality answer, 1.6-1.8x sooner.\n- **Long texts are read in big pieces** (up to 8,192 tokens - pieces of words - at a time), which is why a long\ndocument or code base is read at over 1,000 tokens per second.\n\nWant the full picture? The [details](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#how-it-works) explain every part and its numbers, and the\n[paper](https://github.com/Niko1221/Strata/blob/main/docs/paper/Strata-Paper.pdf) tells the whole story, with the measurements behind it.\n\n- Model: [Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) by the Qwen team; compressed versions by[ISTA-DASLab](https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF) ;[Swift 1.5](https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF) by UkisAI. Their licenses apply\nto the model files.\n- Built with parts of [llama.cpp / ggml](https://github.com/ggml-org/llama.cpp) (MIT). Ideas from[Splash](https://github.com/incoai/splash) ,[ninfer](https://github.com/Neroued/ninfer) and[HyperQwen](https://github.com/syv-ai/HyperQwen) . More in the[details](https://github.com/Niko1221/Strata/blob/main/docs/DETAILS.md#credits-and-licenses) .\n\nStrata is open source under the [MIT License](https://github.com/Niko1221/Strata/blob/main/LICENSE). A few parts carry their own licenses: `third_party/ggml`\n(MIT, llama.cpp / ggml), the web app's font (SIL Open Font License 1.1) and the experimental speed projection's\nvector in `data/experimental-speed-projection` (Qwen Community License 1.0, from the model's activations). The\nmodels are not part of this repository; each model's own license applies to its files.", "url": "https://wpnews.pro/news/strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu", "canonical_source": "https://github.com/Niko1221/Strata", "published_at": "2026-09-28 15:07:39+00:00", "updated_at": "2026-09-28 15:20:34.410131+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "ai-infrastructure", "developer-tools"], "entities": ["Strata", "Qwen3.8-Flash-Next", "Niko1221", "NVIDIA", "RTX 5070", "ISTA-DASLab", "UkisAI", "Swift 1.5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu", "markdown": "https://wpnews.pro/news/strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu.md", "text": "https://wpnews.pro/news/strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu.txt", "jsonld": "https://wpnews.pro/news/strata-qwen3-8-flash-next-125b-moe-on-a-8gb-nvidia-gpu.jsonld"}}