22:33
2026-07-24
gilesthomas.com
large-language-models
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090
Benchmarking Qwen 3.6 35B MoE (3B active) on an RTX 3090 using Unsloth's UD-IQ4_NL_XL quantisation achieved up to 140 tokens per second for generation and over 3,300 tok/s for prompt processing with aโฆ