Qwen 2.5-32B MoE on RTX 3090: Performance Report
Qwen 2.5-32B MoE, a Mixture-of-Experts model activating only about 3B parameters per token, runs efficiently on a single RTX 3090 with 18-22GB VRAM usage and high tokens per second, outperforming stan…