{"slug": "deepseek-v4-flash-0731-latency-numbers-from-nine-providers", "title": "Deepseek V4 Flash 0731 latency numbers from nine providers", "summary": "A one-time snapshot of DeepSeek V4 Flash 0731 latency across nine inference providers found Baseten fastest with 3,980 decode tok/s and 7.72 s total p99, while Azure ran an older checkpoint and Scaleway logged 18 HTTP 429 errors. The test sent 30 concurrent streaming requests per provider across four workloads, totaling 1,080 requests.", "body_md": "# Deepseek V4 Flash 0731 latency numbers from nine providers\n\nDeepSeek V4 Flash 0731 is available from many inference providers. I wanted to see which ones could handle production traffic.\n\nI sent 30 streaming requests concurrently to each provider across four workloads: 120 per provider and 1,080 total. Decode speed is the combined rate of the 30 decode requests.\n\nThis is a one-time snapshot. DeepInfra felt faster earlier, before US traffic ramped up.\n\n## Results\n\n| Provider | Decode tok/s | TTFT p50 / p99 | Total p99 | Cache hit | Errors |\n|---|---|---|---|---|---|\n| Scaleway | 1,650 | 733 / 818 ms | 18.62 s | 80.2% | 18/120 |\n| TensorX | 1,192 | 873 / 1,630 ms | 25.77 s | 73.4% | 1/120 |\n| Fireworks | 2,333 | 1,072 / 2,592 ms | 13.17 s | 17.6% | 0/120 |\n| Baseten | 3,980 |\n774 / 2,979 ms | 7.72 s |\n78.7% | 0/120 |\n| Azure | 596 | 618 / 3,029 ms |\n51.55 s | 57.7% | 0/120 |\n| DeepInfra | 685 | 1,466 / 4,663 ms | 44.85 s | 86.6% | 0/120 |\n| DigitalOcean | 1,280 | 985 / 7,412 ms | 24.00 s | 26.5% | 0/120 |\n| Nebius | 2,542 | 948 / 7,563 ms | 12.08 s | 0.0% | 0/120 |\n| Lyceum | 934 | 8,275 / 9,354 ms | 32.89 s | 96.0% |\n0/120 |\n\nAzure wasn’t running the new DeepSeek V4 Flash 0731 checkpoint, so its results aren’t directly comparable with the others.\n\nScaleway’s 18 failures were HTTP 429s from a token-per-minute quota; successful requests had the best TTFT distribution. TensorX had one transient HTTP 502. Everything else succeeded.", "url": "https://wpnews.pro/news/deepseek-v4-flash-0731-latency-numbers-from-nine-providers", "canonical_source": "https://oskrim.github.io/engineering/2026/08/14/deepseek-v4-flash-providers.html", "published_at": "2026-08-14 00:00:00+00:00", "updated_at": "2026-08-14 12:37:30.238887+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["DeepSeek V4 Flash 0731", "Scaleway", "TensorX", "Fireworks", "Baseten", "Azure", "DeepInfra", "DigitalOcean"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-flash-0731-latency-numbers-from-nine-providers", "markdown": "https://wpnews.pro/news/deepseek-v4-flash-0731-latency-numbers-from-nine-providers.md", "text": "https://wpnews.pro/news/deepseek-v4-flash-0731-latency-numbers-from-nine-providers.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-flash-0731-latency-numbers-from-nine-providers.jsonld"}}