15:50
2026-09-30
insufferable.dev
ai-infrastructure
The AI Race Just Got Awkward
Chinese AI labs' KV cache optimizations, led by DeepSeek's MLA architecture and its DeepSeek-V4.1-Flash release, cut the global KV cache to 890 bytes per token and reduced the cache footprint for longβ¦