cd /news/large-language-models/qwen3-8-27b-at-200-tok-s-peak-on-an-… · home › topics › large-language-models › article
[ARTICLE · art-147900] src=twitter.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max

Zhihao Jia announced the open-sourcing of lithos-metal, a project that runs Qwen3.8-27B at 200+ tokens/s per user peak on a single Apple M5 Max using megakernels and DSpark speculative decoding. The code is available at github.com/lithos-ai/lithos-metal and a technical blog post at lithosai.com, with the claim that it can be tried with any coding agent in one command. A commenter questioned how much of the 200 tok/s peak holds up in a real multi-turn agent loop with tool calls once the KV cache grows, citing Apple Silicon's memory bandwidth limits.

read1 min views1 publishedOct 8, 2026
Qwen3.8-27B at ~200 tok/s peak on an Apple M5 Max
Image: source

Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.

Code: https://t.co/A396OYmw1S
Tech blog: https://t.c… / X

Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀

Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.

Code: https://t.co/A396OYmw1S
Tech blog: https://t.co/bJIlwz9462"

We’re open-sourcing lithos-metal 🚀

Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.

Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…

We’re open-sourcing lithos-metal 🚀

Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.

Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…

nice peak numbers, but apple silicon still runs into the memory bandwidth wall pretty quick once the kv cache grows. curious how much of that 200 tok/s holds up in a real multi-turn agent loop with tool calls instead of pure generation

── more in #large-language-models 4 stories · sorted by recency
── more on @zhihao jia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/qwen3-8-27b-at-200-t…] indexed:0 read:1min 2026-10-08 · —