{"slug": "deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate", "title": "DeepSeek v4.1 Flash at 378 tok/s 99.7% Cache hit rate", "summary": "RunInfra is serving DeepSeek V4.1 Flash (deepseek-ai/DeepSeek-V4.1-Flash) at a promotional $0.10 per 1M input tokens, $0.02 per 1M cached input tokens and $0.43 per 1M output tokens — 25% off standard rates of $0.14, $0.03 and $0.58 — until Oct 13, 2026, 5:44 AM UTC, after which standard rates resume automatically. The model offers a 1,048,576-token context window, OpenAI-compatible chat completions and Anthropic-compatible Messages via POST /v1/messages, with automatic prefix caching on every replica and zero data retention by default. RunInfra reports measured performance of 378 tokens per second and a 99.7% cache hit rate.", "body_md": "# DeepSeek V4.1 Flash\n\n`deepseek-ai/DeepSeek-V4.1-Flash`\nDeepSeek V4.1 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4.1-Flash at $0.10 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.\n\n## Pricing\n\nUSD, pay per token\n\n- per 1M input tokens\n- $0.10\n- per 1M cached input tokens\n- $0.02\n- per 1M output tokens\n- $0.43\n\nNormally $0.14\n\nNormally $0.03\n\nNormally $0.58\n\nInput and output are 25% off until .\n\n05days\n\n22hrs\n\n55min\n\n38sec\n\nEnds in 5 days 22 hours\nStandard rates resume automatically when the window ends.\n\n## Measured performance\n\n## Access\n\nConfirm how your client reaches this model.\n\n- Provider\n- DeepSeek\n- API compatibility\n- OpenAI-compatible chat completions\n- Anthropic compatibility\n- Anthropic-compatible Messages, POST /v1/messages\n- Accepted input\n- Text only\n- Availability\n- Available\n- Data retention\n- [Zero data retention by default. Never used for training.](https://runinfra.ai/docs/security/data-retention)\n\n## Capacity\n\nCheck the limits your workload must fit.\n\n- Context window\n- 1,048,576 tokens\n- Maximum request size\n- 3.5 MB per request\n- Maximum generated output\n- 1,048,576 tokens\n\n## Capabilities\n\nSee which request modes the API supports.\n\n- Tool calling\n- Supported\n- JSON mode\n- Supported\n- Streaming\n- Supported\n\n## View full spec\n\n- Gateway compatibility\n- OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways\n- Prefix caching\n- Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.\n- Cache retention\n- Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.\n- Upstream model\n- [View model](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)", "url": "https://wpnews.pro/news/deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate", "canonical_source": "https://runinfra.ai/inference-api/deepseek-v4-1-flash", "published_at": "2026-10-07 06:33:40+00:00", "updated_at": "2026-10-07 06:49:10.570329+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-products", "ai-tools"], "entities": ["DeepSeek", "DeepSeek V4.1 Flash", "RunInfra", "RunInfra Model APIs", "Hugging Face"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate", "markdown": "https://wpnews.pro/news/deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate.md", "text": "https://wpnews.pro/news/deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-1-flash-at-378-tok-s-99-7-cache-hit-rate.jsonld"}}