{"slug": "new-model-available-deepseek-v4-1-flash", "title": "New Model Available: DeepSeek V4.1 Flash", "summary": "DeepSeek released DeepSeek V4.1 Flash, a 552B-parameter mixture-of-experts model built on a new causal encoder-decoder architecture that natively supports multimodal visual understanding. The model is priced at $0.15-0.3 per million input tokens and $0.6-1.2 per million output tokens, with cache reads at 0.003-0.006 per million tokens, and is positioned for high-throughput, cost-sensitive agentic workloads through reduced KV cache requirements.", "body_md": "DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a new causal encoder-decoder architecture, designed for higher capability, faster reasoning, higher throughput, and lower serving cost. It natively supports multimodal visual understanding and delivers flagship-level intelligence with significantly reduced KV cache requirements, making it well suited for high-throughput and cost-sensitive agentic workloads.\n\n[Back to Models](/models)\n\n## Providers\n\nRoute requests across multiple providers. Copy a provider slug to set your preference.\n\n**$0.15-0.3**\n\n*/ M tokens*\n\n**$0.6-1.2**\n\n*/ M tokens*\n\n*Read:*\n\n**0.003-0.006**/ M tokens\n\n*Write:*\n\n**-**/ M tokens1M1.4s123tps\n\n## Uptime\n\n24hours\nDirect request success rate on AI Gateway and per-provider.\n\n## Throughput\n\n24hours\nP50 throughput on live AI Gateway traffic, in tokens per second (TPS).\n\n## Latency\n\n24hours\nP50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.\n\n## Activity\n\nToken volume and request traffic to this model over time.\n\n## Apps\n\nPublic apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. [View All](/analytics/apps)\n\n## Related Models\n\nMore models from [DeepSeek](/deepseek)", "url": "https://wpnews.pro/news/new-model-available-deepseek-v4-1-flash", "canonical_source": "https://zenmux.ai/deepseek/deepseek-v4.1-flash", "published_at": "2026-09-10 06:23:01+00:00", "updated_at": "2026-09-10 07:21:57.054985+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "generative-ai", "ai-infrastructure", "ai-agents"], "entities": ["DeepSeek", "DeepSeek V4.1 Flash", "AI Gateway"], "alternates": {"html": "https://wpnews.pro/news/new-model-available-deepseek-v4-1-flash", "markdown": "https://wpnews.pro/news/new-model-available-deepseek-v4-1-flash.md", "text": "https://wpnews.pro/news/new-model-available-deepseek-v4-1-flash.txt", "jsonld": "https://wpnews.pro/news/new-model-available-deepseek-v4-1-flash.jsonld"}}