15:33
2026-08-28
github.com
artificial-intelligence
GLM-5.3-Flash on Apple Silicon
WARP, an embeddable inference engine written in C, now runs the full 2.78-trillion-parameter Kimi K3 model on a 64 GB MacBook Pro at about 0.6 tokens per second, and the 313-billion-parameter GLM-5.3-…