DeepSeek V4.1 Flash Specs: KV Cache Compression Explained
DeepSeek AI's model card for DeepSeek V4.1 Flash reports a global KV cache footprint of roughly 890 bytes per token, about a quarter of the 4x-larger footprint of predecessor DeepSeek V4 Flash and a 437x reduction versus…