New Model Available: DeepSeek V4.1 Flash DeepSeek released DeepSeek V4.1 Flash, a 552B-parameter mixture-of-experts model built on a new causal encoder-decoder architecture that natively supports multimodal visual understanding. The model is priced at $0.15-0.3 per million input tokens and $0.6-1.2 per million output tokens, with cache reads at 0.003-0.006 per million tokens, and is positioned for high-throughput, cost-sensitive agentic workloads through reduced KV cache requirements. DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a new causal encoder-decoder architecture, designed for higher capability, faster reasoning, higher throughput, and lower serving cost. It natively supports multimodal visual understanding and delivers flagship-level intelligence with significantly reduced KV cache requirements, making it well suited for high-throughput and cost-sensitive agentic workloads. Back to Models /models Providers Route requests across multiple providers. Copy a provider slug to set your preference. $0.15-0.3 / M tokens $0.6-1.2 / M tokens Read: 0.003-0.006 / M tokens Write: - / M tokens1M1.4s123tps Uptime 24hours Direct request success rate on AI Gateway and per-provider. Throughput 24hours P50 throughput on live AI Gateway traffic, in tokens per second TPS . Latency 24hours P50 time to first token TTFT on live AI Gateway traffic, in milliseconds. Activity Token volume and request traffic to this model over time. Apps Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All /analytics/apps Related Models More models from DeepSeek /deepseek