On a 16 GB RTX 4070 Ti SUPER, Qwen-Image-2.1-Turbo does not fit in bf16 (the pipeline is 32.5 GB). Quantized to NF4 with bitsandbytes it does, and I ran it against the FLUX.1-dev Q6_K + pixel-art LoRA that made my blog covers: same prompts, same seeds, same card, every image checked by eye.
What I measured
Speed: 8.4 s per image for Qwen NF4 (8 steps) against 61 s for FLUX + LoRA (30 steps), median, 1344×768. About 7×. #
Text in the image: 13 of 18 requested texts exact for Qwen NF4, 7 for FLUX. Unquantized Qwen on a 48 GB A6000 got 17. #
Cover prompts: Qwen NF4 drew the requested scene in 11 of 15 covers; FLUX + LoRA in none (6 almost, 9 missing the main action or object). #
Memory: 11.9 GiB over the card's baseline for Qwen NF4 (nvidia-smi), 13.1 GiB for FLUX.
What broke on the way
- An 8-bit (LLM.int8) transformer returns a near-uniform purple field. NF4 works.
- Decoding the VAE in tiles (needed to fit) left faint vertical purple streaks. Fix: make the latents first, drop the transformer and the text encoder from the GPU, then decode the whole frame (0.77 s, 7.1 GiB peak).
- An automatic reader (gemma3:4b) accepted misspelled text, and an external audit caught five of my own verdicts. Every number in the post comes from the corrected image-by-image review.
What I lose: the LoRA's atmosphere. Qwen follows the scene but on a flatter background. I switched anyway: every cover on the blog is now generated by Qwen NF4.
**Read the full measurements, the side-by-side galleries and the code →** [https://efraingaray.com/en/blog/qwen-image-vs-flux/](https://efraingaray.com/en/blog/qwen-image-vs-flux/)
Scripts, prompts and the per-image review: [gist](https://gist.github.com/EfrainGaray/427b5cc1875b888fd57e32bf8b886527).