cd /news/large-language-models/deepseek-v4-1-flash · home topics large-language-models article
[ARTICLE · art-125497] src=twitter.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

DeepSeek v4.1 Flash

DeepSeek released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, now live on the DeepSeek API with native multimodal support under the model name deepseek-flash. The 552B-parameter mixture-of-experts model uses a new Causal Encoder–Decoder architecture with just 8B active parameters for input and 16B for output, and its KV cache requires 1/4 the HBM and 1/8 the SSD storage of the previous generation. DeepSeek retired V4-Flash and V4-Flash-Vision-Exp, temporarily routing deepseek-v4-flash and deepseek-v4-flash-vision-exp to V4.1-Flash, and set off-peak API rates at 50% of peak rates.

read2 min views2 publishedSep 10, 2026
DeepSeek v4.1 Flash
Image: source

DeepSeek on X: "🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6" / X

DeepSeek on X: "🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6"

🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6

🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6

🧠 Asymmetric architecture. More intelligence, less cost. 🔹 552B-parameter MoE. 🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. 🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead ofShow more

💾 Smaller KV cache. Bigger savings. Compared with the previous generation, V4.1-Flash’s KV cache needs just: 🔹 1/4 the HBM 🔹 1/8 the SSD storage Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly. 3/6

⚡ V4.1-Flash is now live on the DeepSeek API with native multimodal support. Set your model to deepseek-flash. 🔹 V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. 🔹 Tests byShow more

💰 More efficient architecture. Lower API prices. V4.1-Flash lets us serve more users at a lower cost. We’re passing the savings on to you. 🔹 Peak/off-peak pricing continues to balance demand. 🔹 Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak toShow more

🌐 Supporting open source. Expanding deployment options. We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk. 🔹 Model:Show more

Quick one: DeepSeek V4.1 Flash lands on WorkBuddy/CodeBuddy, co-launched exclusively in China, discounted for two weeks. TokenHub, ima, and Marvis are live day 0 as well. Well played, @deepseek_ai 🫡 V4.1 Flash is a really solid model, major gains across text and agentShow more

── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-v4-1-flash] indexed:0 read:2min 2026-09-10 ·