# Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max

> Source: <https://twitter.com/JiaZhihao/status/2108249739414147259>
> Published: 2026-10-08 22:15:15+00:00

Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.
Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: https://t.co/A396OYmw1S
Tech blog: https://t.c… / X

Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.
Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: https://t.co/A396OYmw1S
Tech blog: https://t.co/bJIlwz9462"

We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.
Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…

We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max.
Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…

nice peak numbers, but apple silicon still runs into the memory bandwidth wall pretty quick once the kv cache grows. curious how much of that 200 tok/s holds up in a real multi-turn agent loop with tool calls instead of pure generation
