cd /news/machine-learning/show-hn-lora-over-gguf-train-qwen3-8… · home › topics › machine-learning › article
[ARTICLE · art-148826] src=github.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

Show HN: LoRA over GGUF – Train Qwen3.8-Flash-Next in 40G VRAM

A developer demonstrated training Qwen3.8-Flash-Next (125B-A6B plus 51B engram) in 40 GiB VRAM without CPU offload using LoRA over GGUF, reaching 200 token/s on Strix Halo. The method relies on HuggingFace Transformers 5.18's new GGUF support, which handles MoE and sparse-attention models that bnb 4-bit does not, and the same setup trains DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM at 100 token/s.

read1 min views1 publishedOct 10, 2026

Open weight AI is like open source software. Users not only run the weights, but also modify the weights. It matters to develop training framework for local hardware.

The method to train large AI models on local hardware is generally called QLoRA (Low Rank Adaptation over Quantized base model). In the last few years it's usually done with HuggingFace Transformers (which is the basis of training frameworks such as Unsloth and Axolotl) and bnb 4-bit base model. However, bnb does not yet support MoE models, so the local training of recent MoE models seemed stall for some time.

Since Transformers 5.18, it's started to support GGUF, and it supports recent models with MoE and sparse attentions (and it's not slow, already faster than llama.cpp on Mac). GGUF is a versatile container format. It can be smaller than 4 bpw with surprisingly good quantization quality, so there are new possibilities for local training.

I've shown that we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM without CPU offload. I've optimized it on Strix Halo and it trains at 200 token/s. There is still room to optimize, compared to > 1600 token/s prompt processing we've achieved, and the common sense that LoRA training (with gradient checkpointing) takes 4-5x work of prompt processing. CPU/disk offload (like Strata) and multi-GPU also need more work that I'm not currently focusing on.

On Strix Halo we can also train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM at 100 token/s, but I think it's less practical than Qwen3.8FN for local use.

Comments URL: [https://news.ycombinator.com/item?id=50034625](https://news.ycombinator.com/item?id=50034625)

Points: 1

── more in #machine-learning 4 stories · sorted by recency
── more on @qwen3.8-flash-next 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-lora-over-gg…] indexed:0 read:1min 2026-10-10 · —