Trying to improve output speed with MLX models on LM Studio An LM Studio user running version 0.4.25+1 on a MacBook Pro M5 Max with 128 GB of RAM reports getting 15-17 tokens per second from the mlx-community/Qwen3.8-27B-8bit model and cannot enable speculative decoding because LM Studio returns "No compatible draft models found for your current model selection" after downloading mlx-community/Qwen3.8-27B-MTP-8bit as a draft model. The user is running the LM Studio MLX (Apple M5) v1.11.0 engine and is asking for an alternative draft model or other ways to improve output speed. Hello, I’ve been using LM Studio on my Macbook Pro M5 Max 128 GB with mlx-community/qwen3.8-27b https://huggingface.co/mlx-community/Qwen3.8-27B-8bit without issues, so I’m looking for ways to improve output speed right now I’m getting 15-17 t/s and found about speculative decoding. Looking around I found mlx-community/Qwen3.8-27B-MTP-8bit https://huggingface.co/mlx-community/Qwen3.8-27B-MTP-8bit is a draft model for that, so I downloaded it, but when I try to enable speculative decoding, LM Studio says “No compatible draft models found for your current model selection”. My LM Studio version is 0.4.25+1 and the MLX engine I’m using is LM Studio MLX Apple M5 v1.11.0. If anyone can suggest another draft model or any other way to improve output speed I’d really appreciate it. Thanks in advance.