Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max Zhihao Jia announced the open-sourcing of lithos-metal, a project that runs Qwen3.8-27B at 200+ tokens/s per user peak on a single Apple M5 Max using megakernels and DSpark speculative decoding. The code is available at github.com/lithos-ai/lithos-metal and a technical blog post at lithosai.com, with the claim that it can be tried with any coding agent in one command. A commenter questioned how much of the 200 tok/s peak holds up in a real multi-turn agent loop with tool calls once the KV cache grows, citing Apple Silicon's memory bandwidth limits. Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command. Code: https://t.co/A396OYmw1S Tech blog: https://t.c… / X Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command. Code: https://t.co/A396OYmw1S Tech blog: https://t.co/bJIlwz9462" We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command. Code: github.com/lithos-ai/lith… Tech blog: lithosai.com/blog/lithos-me… We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command. Code: github.com/lithos-ai/lith… Tech blog: lithosai.com/blog/lithos-me… nice peak numbers, but apple silicon still runs into the memory bandwidth wall pretty quick once the kv cache grows. curious how much of that 200 tok/s holds up in a real multi-turn agent loop with tool calls instead of pure generation