# New Model Available: DeepSeek V4.1 Flash

> Source: <https://zenmux.ai/deepseek/deepseek-v4.1-flash>
> Published: 2026-09-10 06:23:01+00:00

DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a new causal encoder-decoder architecture, designed for higher capability, faster reasoning, higher throughput, and lower serving cost. It natively supports multimodal visual understanding and delivers flagship-level intelligence with significantly reduced KV cache requirements, making it well suited for high-throughput and cost-sensitive agentic workloads.

[Back to Models](/models)

## Providers

Route requests across multiple providers. Copy a provider slug to set your preference.

**$0.15-0.3**

*/ M tokens*

**$0.6-1.2**

*/ M tokens*

*Read:*

**0.003-0.006**/ M tokens

*Write:*

**-**/ M tokens1M1.4s123tps

## Uptime

24hours
Direct request success rate on AI Gateway and per-provider.

## Throughput

24hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS).

## Latency

24hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.

## Activity

Token volume and request traffic to this model over time.

## Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. [View All](/analytics/apps)

## Related Models

More models from [DeepSeek](/deepseek)
