Smaller, faster, safer: running Kimi and GLM at scale
Cloudflare's Workers AI has implemented three optimizations—quantizing the KV cache, compressing model weights, and protecting the shared cache—to run Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 more effi…