New Deepseek model V4.1-Flash cuts memory needs for AI agents Deepseek released V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor's requirement, with only 16 billion parameters active per token. On the DeepSWE coding benchmark, V4.1-Flash narrowly beats Opus 5 and GPT-5.6 Sol, and the model ships under the MIT license targeting cheaper AI agents. Deepseek releases V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents. The article New Deepseek model V4.1-Flash cuts memory needs for AI agents https://the-decoder.com/new-deepseek-model-v4-1-flash-cuts-memory-needs-for-ai-agents/ appeared first on The Decoder https://the-decoder.com .