cd /news/large-language-models/running-qwen-3-8-flash-next-125b-on-… · home › topics › large-language-models › article
[ARTICLE · art-145319] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Running Qwen 3.8 Flash Next (125B) on a RTX 4090 – 100 T/s on a Desktop

A developer known as Niko1221 has published Strata, a repository showing that the 125B-parameter Qwen 3.8 Flash Next model can run on a single consumer RTX 4090 at roughly 100 tokens per second. The setup combines 4-bit bitsandbytes quantization, torch.compile, and a custom CUDA flash-attention kernel to fit the model into 24GB of VRAM, with a sample 128-token generation completing in about 0.9 seconds. The project lowers the hardware barrier for running large models locally, eliminating per-token cloud inference costs.

by read4 min views2 publishedOct 5, 2026

Niko1221’s Strata repo shows that the new Qwen 3.8 Flash Next (125B) model can be run on a consumer‑grade RTX 4090 at an impressive 100 trillion tokens per second (T/s). In practice, that means you can generate text at near‑real‑time speed without a multi‑GPU server or a cloud‑based inference endpoint. The repo ships a set of scripts, quantisation tricks, and a torch.compile‑friendly pipeline that squeezes every ounce of performance out of the 24 GB VRAM card.

Historically, a 125 B‑parameter model was only accessible behind expensive API contracts or on multi‑node GPU clusters. By proving that a single RTX 4090 can handle it, the barrier to entry drops dramatically. Small startups, indie developers, and research teams can now experiment locally, iterate faster, and keep data in‑house for compliance reasons.

Running inference on‑premises eliminates per‑token cloud costs (often $0.0001‑$0.0002 per token) and reduces latency to sub‑second levels for most prompts. That opens up new use‑cases: real‑time assistants, low‑latency code completion, or edge‑ish deployments where a single workstation is the inference node.

From an infra perspective, the trick is not just the model size but how it’s quantised and compiled. Strata uses 4‑bit bitsandbytes + torch.compile + a custom CUDA kernel for flash‑attention. Understanding those pieces lets you apply the same pattern to other massive models (LLaMA‑3, Gemma‑2, etc.) and build reusable pipelines for your own products.

Below is a practical, end‑to‑end walkthrough that I used on a fresh Ubuntu 22.04 machine with the latest NVIDIA driver (560.xx) and CUDA 12.3.

sudo apt update && sudo apt install -y git python3.11 python3.11-venv build-essential

python3.11 -m venv qwen-env
source qwen-env/bin/activate

pip install --upgrade pip
pip install torch==2.3.0+cu123 torchvision==0.18.0+cu123 \
    -f https://download.pytorch.org/whl/torch_stable.html

pip install bitsandbytes==0.43.1

pip install transformers accelerate huggingface_hub

git clone https://github.com/Niko1221/Strata.git
cd Strata
pip install -e .   # installs the tiny helper package
huggingface-cli login   # you need a token for private repos if applicable

python -c "from transformers import AutoModelForCausalLM, AutoTokenizer; \
    AutoModelForCausalLM.from_pretrained('Qwen/Qwen3.8-Flash-Next-125B-4bit', \
    torch_dtype='auto', device_map='auto')"

Strata ships a thin wrapper run_qwen.py. I added a tiny prompt loop to illustrate latency.

python run_qwen.py \
  --model Qwen/Qwen3.8-Flash-Next-125B-4bit \
  --max_new_tokens 128 \
  --temperature 0.7

Sample output (≈0.9 s for 128 tokens on RTX 4090):

User: Explain why the sky is blue in two sentences.
AI: The sky appears blue because molecules in Earth’s atmosphere scatter shorter blue wavelengths of sunlight more than longer red wavelengths. This scattering, known as Rayleigh scattering, sends a lot of blue light toward our eyes.

If you already have a FastAPI or Flask micro‑service, you can reuse the same AutoModelForCausalLM object.

from fastapi import FastAPI, Body
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

app = FastAPI()
model_name = "Qwen/Qwen3.8-Flash-Next-125B-4bit"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype=torch.float16,
    device_map="auto",
    load_in_4bit=True,
)

@app.post("/generate")
async def generate(prompt: str = Body(..., embed=True)):
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
    with torch.cuda.amp.autocast():
        output = model.generate(**inputs, max_new_tokens=150, temperature=0.8)
    return {"response": tokenizer.decode(output[0], skip_special_tokens=True)}

Deploy the service with uvicorn app:app --host 0.0.0.0 --port 8000 and you have a locally hosted 125 B LLM ready for internal tools, chat‑bots, or batch‑processing pipelines.

Running a 125 B model on a single RTX 4090 feels less like a gimmick and more like a new baseline for AI infra. A few observations from my side:

bitsandbytes cuts VRAM usage to ~20 GB while preserving >90 % of the original quality for most conversational tasks. The trade‑off is negligible for many internal applications.torch.compile (beta) together with the flash‑attention kernel gives us the 100 T/s claim. Without it, latency balloons to 2‑3 s per 128 tokens. In short, the Strata approach proves that massive LLMs are no longer the exclusive domain of hyperscale clouds. By leveraging 4‑bit quantisation, torch‑compile, and flash‑attention, a single RTX 4090 can give you a production‑grade inference engine at a fraction of the cost. As an AI infrastructure engineer, I see this as a turning point: the next wave of AI products will be built on local giant models, and the tooling we adopt today will become the foundation of that wave.

── more in #large-language-models 4 stories · sorted by recency
── more on @niko1221 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/running-qwen-3-8-fla…] indexed:0 read:4min 2026-10-05 · —