Show HN: LoRA over GGUF – Train Qwen3.8-Flash-Next in 40G VRAM
A developer demonstrated training Qwen3.8-Flash-Next (125B-A6B plus 51B engram) in 40 GiB VRAM without CPU offload using LoRA over GGUF, reaching 200 token/s on Strix Halo. The method relies on Huggin…