Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀 Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: https://t.co/A396OYmw1S
Tech blog: https://t.c… / X
Zhihao Jia on X: "We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: https://t.co/A396OYmw1S
Tech blog: https://t.co/bJIlwz9462"
We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…
We’re open-sourcing lithos-metal 🚀
Megakernels + DSpark speculative decoding run Qwen3.8-27B at 200+ tokens/s/user peak on one @Apple M5 Max. Ultra-fast inference on your laptop. Try it with any coding agent in one command.
Code: github.com/lithos-ai/lith…
Tech blog: lithosai.com/blog/lithos-me…
nice peak numbers, but apple silicon still runs into the memory bandwidth wall pretty quick once the kv cache grows. curious how much of that 200 tok/s holds up in a real multi-turn agent loop with tool calls instead of pure generation