DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a new causal encoder-decoder architecture, designed for higher capability, faster reasoning, higher throughput, and lower serving cost. It natively supports multimodal visual understanding and delivers flagship-level intelligence with significantly reduced KV cache requirements, making it well suited for high-throughput and cost-sensitive agentic workloads.
Providers #
Route requests across multiple providers. Copy a provider slug to set your preference.
$0.15-0.3
/ M tokens $0.6-1.2
/ M tokens Read:
0.003-0.006/ M tokens
Write:
-/ M tokens1M1.4s123tps
Uptime #
24hours Direct request success rate on AI Gateway and per-provider.
Throughput #
24hours P50 throughput on live AI Gateway traffic, in tokens per second (TPS).
Latency #
24hours P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.
Activity #
Token volume and request traffic to this model over time.
Apps #
Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All
Related Models #
More models from DeepSeek