cd /news/large-language-models/deepseek-v4-1-flash-at-378-tok-s-99-… · home › topics › large-language-models › article
[ARTICLE · art-146641] src=runinfra.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

DeepSeek v4.1 Flash at 378 tok/s 99.7% Cache hit rate

RunInfra is serving DeepSeek V4.1 Flash (deepseek-ai/DeepSeek-V4.1-Flash) at a promotional $0.10 per 1M input tokens, $0.02 per 1M cached input tokens and $0.43 per 1M output tokens — 25% off standard rates of $0.14, $0.03 and $0.58 — until Oct 13, 2026, 5:44 AM UTC, after which standard rates resume automatically. The model offers a 1,048,576-token context window, OpenAI-compatible chat completions and Anthropic-compatible Messages via POST /v1/messages, with automatic prefix caching on every replica and zero data retention by default. RunInfra reports measured performance of 378 tokens per second and a 99.7% cache hit rate.

read2 min views1 publishedOct 7, 2026
DeepSeek v4.1 Flash at 378 tok/s 99.7% Cache hit rate
Image: source

deepseek-ai/DeepSeek-V4.1-Flash DeepSeek V4.1 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4.1-Flash at $0.10 per 1M input tokens until Oct 13, 2026, 5:44 AM UTC, with standard rates resuming automatically. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions and Anthropic-compatible Messages, POST /v1/messages.

Pricing #

USD, pay per token

  • per 1M input tokens
  • $0.10
  • per 1M cached input tokens
  • $0.02
  • per 1M output tokens
  • $0.43

Normally $0.14

Normally $0.03

Normally $0.58

Input and output are 25% off until .

05days

22hrs

55min

38sec

Ends in 5 days 22 hours Standard rates resume automatically when the window ends.

Measured performance #

Access #

Confirm how your client reaches this model.

Capacity #

Check the limits your workload must fit.

  • Context window
  • 1,048,576 tokens
  • Maximum request size
  • 3.5 MB per request
  • Maximum generated output
  • 1,048,576 tokens

Capabilities #

See which request modes the API supports.

  • Tool calling
  • Supported
  • JSON mode
  • Supported
  • Streaming
  • Supported

View full spec #

  • Gateway compatibility
  • OpenAI- and Anthropic-compatible APIs for compatible clients, tools, and gateways
  • Prefix caching
  • Automatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation or, failing that, the workspace. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts.
  • Cache retention
  • Prefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort.
  • Upstream model
  • View model
── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-v4-1-flash-…] indexed:0 read:2min 2026-10-07 · —