DeepSeek v4.1 Flash at 378 tok/s 99.7% Cache hit rate RunInfra is serving DeepSeek V4.1 Flash (deepseek-ai/DeepSeek-V4.1-Flash) at a promotional $0.10 per 1M input tokens, $0.02 per 1M cached input tokens and $0.43 per 1M output tokens — 25% off standard rates of $0.14, $0.03 and $0.58 — until Oct 13, 2026, 5:44 AM UTC, after which standard rates resume automatically. The model offers a 1,048,576-token context window, OpenAI-compatible chat completions and Anthropic-compatible Messages via POST /v1/messages, with automatic prefix caching on every replica and zero data retention by default. RunInfra reports measured performance of 378 tokens per second and a 99.7% cache hit rate. DeepSeek V4.1 Flash deepseek-ai/DeepSeek-V4.1-Flash DeepSeek V4.1 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4.1-Flash at $0.10 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages. Pricing USD, pay per token - per 1M input tokens - $0.10 - per 1M cached input tokens - $0.02 - per 1M output tokens - $0.43 Normally $0.14 Normally $0.03 Normally $0.58 Input and output are 25% off until . 05days 22hrs 55min 38sec Ends in 5 days 22 hours Standard rates resume automatically when the window ends. Measured performance Access Confirm how your client reaches this model. - Provider - DeepSeek - API compatibility - OpenAI-compatible chat completions - Anthropic compatibility - Anthropic-compatible Messages, POST /v1/messages - Accepted input - Text only - Availability - Available - Data retention - Zero data retention by default. Never used for training. https://runinfra.ai/docs/security/data-retention Capacity Check the limits your workload must fit. - Context window - 1,048,576 tokens - Maximum request size - 3.5 MB per request - Maximum generated output - 1,048,576 tokens Capabilities See which request modes the API supports. - Tool calling - Supported - JSON mode - Supported - Streaming - Supported View full spec - Gateway compatibility - OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways - Prefix caching - Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt cache key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts. - Cache retention - Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort. - Upstream model - View model https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash