You Got mlx-serve'd!
Mark Murphy reports that mlx-serve, an Ollama-like inference server with a GUI, delivers a 4x speed increase for running Qwen 3.8 on his 64GB M2 Ultra Mac Studio compared to his previous Ollama setup,β¦
Mark Murphy reports that mlx-serve, an Ollama-like inference server with a GUI, delivers a 4x speed increase for running Qwen 3.8 on his 64GB M2 Ultra Mac Studio compared to his previous Ollama setup,β¦
Qwen 3.8, a 27B-parameter open-weights model from Alibaba, delivers coding output in the Claude Haiku-to-Sonnet range but is 15-20 times slower than Claude Sonnet, according to tests by a developer onβ¦
Meta's open-weights Muse-Glimmer model, run locally via Ollama on a 64GB M2 Ultra Mac Studio, took more than twice as long as Qwen 3.6 to complete a documentation-review prompt and produced a small frβ¦
Local models such as Qwen 3.6 and Gemma 4 are close to being useful for Kotlin Multiplatform code generation but still fall short for many developers, according to Mark Murphy's experiments. Murphy suβ¦