{"slug": "qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max", "title": "Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max", "summary": "Zhihao Jia announced the open-sourcing of lithos-metal, a project that runs Qwen3.8-27B at 200+ tokens/s per user peak on a single Apple M5 Max using megakernels and DSpark speculative decoding. The code is available at github.com/lithos-ai/lithos-metal and a technical blog post at lithosai.com, with the claim that it can be tried with any coding agent in one command. A commenter questioned how much of the 200 tok/s peak holds up in a real multi-turn agent loop with tool calls once the KV cache grows, citing Apple Silicon's memory bandwidth limits.", "body_md": "Zhihao Jia on X: \"We’re open-sourcing lithos-metal 🚀\nMegakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.\nUltra-fast inference on your laptop. Try it with any coding agent in one command.\nCode: https://t.co/A396OYmw1S\nTech blog: https://t.c… / X\n\nZhihao Jia on X: \"We’re open-sourcing lithos-metal 🚀\nMegakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.\nUltra-fast inference on your laptop. Try it with any coding agent in one command.\nCode: https://t.co/A396OYmw1S\nTech blog: https://t.co/bJIlwz9462\"\n\nWe’re open-sourcing lithos-metal 🚀\nMegakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.\nUltra-fast inference on your laptop. Try it with any coding agent in one command.\nCode: github.com/lithos-ai/lith…\nTech blog: lithosai.com/blog/lithos-me…\n\nWe’re open-sourcing lithos-metal 🚀\nMegakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.\nUltra-fast inference on your laptop. Try it with any coding agent in one command.\nCode: github.com/lithos-ai/lith…\nTech blog: lithosai.com/blog/lithos-me…\n\nnice peak numbers, but apple silicon still runs into the memory bandwidth wall pretty quick once the kv cache grows. curious how much of that 200 tok/s holds up in a real multi-turn agent loop with tool calls instead of pure generation", "url": "https://wpnews.pro/news/qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max", "canonical_source": "https://twitter.com/JiaZhihao/status/2108249739414147259", "published_at": "2026-10-08 22:15:15+00:00", "updated_at": "2026-10-08 22:47:47.012600+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools", "developer-tools"], "entities": ["Zhihao Jia", "lithos-metal", "Qwen3.8-27B", "Apple M5 Max", "DSpark", "lithos-ai", "Apple"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max", "markdown": "https://wpnews.pro/news/qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max.md", "text": "https://wpnews.pro/news/qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max.txt", "jsonld": "https://wpnews.pro/news/qwen3-8-27b-at-200-tok-s-peak-on-an-apple-m5-max.jsonld"}}