07:31
2026-09-10
marktechpost.com
large-language-models
DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reuse
DeepSeek AI released DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B backbone parameters, 196B additional Engram parameters, and a 1M-token context window, cutting global KV cacheβ¦