DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression DeepSeek released DeepSeek-V4.1-Flash, a model that targets KV cache compression to reduce the memory and storage burden of long-context, input-heavy agent workloads. DeepSeek said prefill remains computationally expensive and large KV caches continue to strain HBM and SSD capacity and data transfer. The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transf