12:40
2026-09-10
the-decoder.com
large-language-models
New Deepseek model V4.1-Flash cuts memory needs for AI agents
Deepseek released V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor's requirement, with only 16 billion parameters active per token. β¦