# DeepSeek v4.1 Flash at 378 tok/s 99.7% Cache hit rate

> Source: <https://runinfra.ai/inference-api/deepseek-v4-1-flash>
> Published: 2026-10-07 06:33:40+00:00

# DeepSeek V4.1 Flash

`deepseek-ai/DeepSeek-V4.1-Flash`
DeepSeek V4.1 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4.1-Flash at $0.10 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.

## Pricing

USD, pay per token

- per 1M input tokens
- $0.10
- per 1M cached input tokens
- $0.02
- per 1M output tokens
- $0.43

Normally $0.14

Normally $0.03

Normally $0.58

Input and output are 25% off until .

05days

22hrs

55min

38sec

Ends in 5 days 22 hours
Standard rates resume automatically when the window ends.

## Measured performance

## Access

Confirm how your client reaches this model.

- Provider
- DeepSeek
- API compatibility
- OpenAI-compatible chat completions
- Anthropic compatibility
- Anthropic-compatible Messages, POST /v1/messages
- Accepted input
- Text only
- Availability
- Available
- Data retention
- [Zero data retention by default. Never used for training.](https://runinfra.ai/docs/security/data-retention)

## Capacity

Check the limits your workload must fit.

- Context window
- 1,048,576 tokens
- Maximum request size
- 3.5 MB per request
- Maximum generated output
- 1,048,576 tokens

## Capabilities

See which request modes the API supports.

- Tool calling
- Supported
- JSON mode
- Supported
- Streaming
- Supported

## View full spec

- Gateway compatibility
- OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
- Prefix caching
- Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
- Cache retention
- Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
- Upstream model
- [View model](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
