15:31
2026-09-23
discuss.huggingface.co
large-language-models
Trying to improve output speed with MLX models on LM Studio
An LM Studio user running version 0.4.25+1 on a MacBook Pro M5 Max with 128 GB of RAM reports getting 15-17 tokens per second from the mlx-community/Qwen3.8-27B-8bit model and cannot enable speculativ…