{"slug": "splash-engine-the-fastest-local-qwen3-8-on-apple-silicon", "title": "Splash Engine - the fastest local Qwen3.8 on Apple Silicon", "summary": "Inco AI released Splash, an open-source inference engine for running Qwen3.6-35B-A3B and Qwen3.8-27B locally on Apple silicon, which in Inco's tests on a 48 GB M5 Pro delivered 74 tokens per second on short prompts and 54 at 32K context on Qwen3.8-27B — roughly twice the decode speed of the next-fastest engine measured. With four concurrent requests on short prompts, Splash reached 170 tokens per second combined, 3.9× the next-fastest engine in the comparison. Splash is available through LM Studio Bionic 1.1.5 or newer and requires an M3-or-newer Mac running macOS 26.4 or later with at least 36 GB of unified memory, with Inco recommending 48 GB or more.", "body_md": "# Splash Engine - the fastest local Qwen3.8 on Apple Silicon\n\n## What is Splash Engine?\n\nSplash is an open-source inference engine from [Inco AI](https://inco.ai/) for running language models locally on Apple silicon. It is optimized specifically for Qwen3.6-35B-A3B and Qwen3.8-27B. The engine provides GPU kernels and a memory plan tailored to each supported model. Each model ships with a dedicated DFlash 2 draft model for speculative decoding, which improves generation speed.\n\nIn Inco's tests on a 48 GB M5 Pro, Splash delivered roughly twice the decode speed of the next-fastest engine they measured on Qwen3.8-27B: 74 tokens per second on short prompts and 54 at 32K context. With four concurrent requests on short prompts, its combined throughput reached 170 tokens per second—3.9× the next-fastest engine in their comparison. Read more about Splash in Inco's [blog post](https://inco.ai/blog/splash).\n\n## Use it in LM Studio Bionic\n\nDownload and install [LM Studio Bionic](https://lmstudio.ai) 1.1.5 or newer, then open the app. Splash requires an M3-or-newer Mac running macOS 26.4 or later with at least 36 GB of unified memory; Inco recommends 48 GB or more.\n\nNavigate to **Settings > Runtime**. Under **Experimental backends**, click **Download** next to **Splash (Metal)** to install the engine.\n\nDownload the Splash engine from Settings > Runtime.\n\nThen go to **Settings > Explore**, paste one of the following Hugging Face links into the search bar, select the model, and click **Download**:\n\nOnce the download finishes, start a new session and select the model from the local model picker.", "url": "https://wpnews.pro/news/splash-engine-the-fastest-local-qwen3-8-on-apple-silicon", "canonical_source": "https://lmstudio.ai/blog/splash-engine", "published_at": "2026-09-18 00:00:00+00:00", "updated_at": "2026-09-19 00:53:57.989214+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "ai-products"], "entities": ["Inco AI", "Splash Engine", "Qwen3.8-27B", "Qwen3.6-35B-A3B", "LM Studio Bionic", "Apple Silicon", "DFlash 2", "Hugging Face"], "alternates": {"html": "https://wpnews.pro/news/splash-engine-the-fastest-local-qwen3-8-on-apple-silicon", "markdown": "https://wpnews.pro/news/splash-engine-the-fastest-local-qwen3-8-on-apple-silicon.md", "text": "https://wpnews.pro/news/splash-engine-the-fastest-local-qwen3-8-on-apple-silicon.txt", "jsonld": "https://wpnews.pro/news/splash-engine-the-fastest-local-qwen3-8-on-apple-silicon.jsonld"}}