{"slug": "deepseek-v4-1-flash", "title": "DeepSeek v4.1 Flash", "summary": "DeepSeek released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, now live on the DeepSeek API with native multimodal support under the model name deepseek-flash. The 552B-parameter mixture-of-experts model uses a new Causal Encoder–Decoder architecture with just 8B active parameters for input and 16B for output, and its KV cache requires 1/4 the HBM and 1/8 the SSD storage of the previous generation. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, temporarily routing deepseek-v4-flash and deepseek-v4-flash-vision-exp to V4.1-Flash, and set off-peak API rates at 50% of peak rates.", "body_md": "DeepSeek on X: \"🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.\n🔹 Introducing the smallest model in our new architecture family, with native visual understanding.\n🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.\n1/6\" / X\n\nDeepSeek on X: \"🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.\n🔹 Introducing the smallest model in our new architecture family, with native visual understanding.\n🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.\n1/6\"\n\n🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.\n🔹 Introducing the smallest model in our new architecture family, with native visual understanding.\n🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.\n1/6\n\n🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.\n🔹 Introducing the smallest model in our new architecture family, with native visual understanding.\n🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.\n1/6\n\n🧠 Asymmetric architecture. More intelligence, less cost.\n🔹 552B-parameter MoE.\n🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.\n🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead ofShow more\n\n💾 Smaller KV cache. Bigger savings.\nCompared with the previous generation, V4.1-Flash’s KV cache needs just:\n🔹 1/4 the HBM\n🔹 1/8 the SSD storage\nCache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.\n3/6\n\n⚡ V4.1-Flash is now live on the DeepSeek API with native multimodal support.\nSet your model to deepseek-flash.\n🔹 V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.\n🔹 Tests byShow more\n\n💰 More efficient architecture. Lower API prices.\nV4.1-Flash lets us serve more users at a lower cost. We’re passing the savings on to you.\n🔹 Peak/off-peak pricing continues to balance demand.\n🔹 Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak toShow more\n\n🌐 Supporting open source. Expanding deployment options.\nWe’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.\nPlanning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk.\n🔹 Model:Show more\n\nQuick one: DeepSeek V4.1 Flash lands on WorkBuddy/CodeBuddy, co-launched exclusively in China, discounted for two weeks.\nTokenHub, ima, and Marvis are live day 0 as well. Well played, @deepseek_ai 🫡 V4.1 Flash is a really solid model, major gains across text and agentShow more", "url": "https://wpnews.pro/news/deepseek-v4-1-flash", "canonical_source": "https://twitter.com/deepseek_ai/status/2097930608790167907", "published_at": "2026-09-10 06:11:05+00:00", "updated_at": "2026-09-10 06:22:48.504348+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "artificial-intelligence", "ai-infrastructure", "ai-agents"], "entities": ["DeepSeek", "DeepSeek-V4.1-Flash", "DeepSeek API", "V4-Flash", "V4-Flash-Vision-Exp", "WorkBuddy", "CodeBuddy", "TokenHub"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-1-flash", "markdown": "https://wpnews.pro/news/deepseek-v4-1-flash.md", "text": "https://wpnews.pro/news/deepseek-v4-1-flash.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-1-flash.jsonld"}}