{"slug": "trying-to-improve-output-speed-with-mlx-models-on-lm-studio", "title": "Trying to improve output speed with MLX models on LM Studio", "summary": "An LM Studio user running version 0.4.25+1 on a MacBook Pro M5 Max with 128 GB of RAM reports getting 15-17 tokens per second from the mlx-community/Qwen3.8-27B-8bit model and cannot enable speculative decoding because LM Studio returns \"No compatible draft models found for your current model selection\" after downloading mlx-community/Qwen3.8-27B-MTP-8bit as a draft model. The user is running the LM Studio MLX (Apple M5) v1.11.0 engine and is asking for an alternative draft model or other ways to improve output speed.", "body_md": "Hello,\n\nI’ve been using LM Studio on my Macbook Pro M5 Max (128 GB) with [mlx-community/qwen3.8-27b](https://huggingface.co/mlx-community/Qwen3.8-27B-8bit) without issues, so I’m looking for ways to improve output speed (right now I’m getting 15-17 t/s) and found about speculative decoding. Looking around I found [mlx-community/Qwen3.8-27B-MTP-8bit](https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-8bit) is a draft model for that, so I downloaded it, but when I try to enable speculative decoding, LM Studio says “No compatible draft models found for your current model selection”.\n\nMy LM Studio version is 0.4.25+1 and the MLX engine I’m using is LM Studio MLX (Apple M5) v1.11.0.\n\nIf anyone can suggest another draft model or any other way to improve output speed I’d really appreciate it.\n\nThanks in advance.", "url": "https://wpnews.pro/news/trying-to-improve-output-speed-with-mlx-models-on-lm-studio", "canonical_source": "https://discuss.huggingface.co/t/trying-to-improve-output-speed-with-mlx-models-on-lm-studio/180696#post_1", "published_at": "2026-09-23 15:31:44+00:00", "updated_at": "2026-09-23 15:58:43.016829+00:00", "lang": "en", "topics": ["large-language-models", "ai-tools", "mlops"], "entities": ["LM Studio", "Apple", "MacBook Pro M5 Max", "mlx-community/Qwen3.8-27B-8bit", "mlx-community/Qwen3.8-27B-MTP-8bit", "LM Studio MLX (Apple M5) v1.11.0", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/trying-to-improve-output-speed-with-mlx-models-on-lm-studio", "markdown": "https://wpnews.pro/news/trying-to-improve-output-speed-with-mlx-models-on-lm-studio.md", "text": "https://wpnews.pro/news/trying-to-improve-output-speed-with-mlx-models-on-lm-studio.txt", "jsonld": "https://wpnews.pro/news/trying-to-improve-output-speed-with-mlx-models-on-lm-studio.jsonld"}}