You Got mlx-serve'd!
Mark Murphy reports that mlx-serve, an Ollama-like inference server with a GUI, delivers a 4x speed increase for running Qwen 3.8 on his 64GB M2 Ultra Mac Studio compared to his previous Ollama setup,…
Mark Murphy reports that mlx-serve, an Ollama-like inference server with a GUI, delivers a 4x speed increase for running Qwen 3.8 on his 64GB M2 Ultra Mac Studio compared to his previous Ollama setup,…
NVIDIA's June 2026 DGX Spark update introduces automated four-node clustering via Cluster Assistant, enabling local inference of models up to 700B parameters. The update also delivers a 2.6x throughpu…
Google released new Gemma 4 checkpoints optimized with Quantization-Aware Training (QAT) to reduce model memory footprint for local deployment on edge devices and consumer GPUs. The QAT process minimi…
JetBrains released Mellum2, a 12-billion-parameter Mixture-of-Experts model with 2.5 billion active parameters per token, under the Apache 2.0 license. The model is specialized for software engineerin…