DEEPSEEK LATENCY BENCHMARK · Sept 2026
We tested both from six cities. The faster option depends heavily on where you call from.
What the data shows #
- Finding 1 ### Location matters more than the routeDirect was faster from Singapore and Mumbai. OpenRouter was often faster from Montreal.
- Singapore and Mumbai: direct 621–664 ms, routes 782–1,348 ms
- Montreal: routes 422–692 ms, direct 732 ms See the chart by city →
- Finding 2 ### DeepSeek's direct API was more consistentDeepSeek's API varied less between locations: its slowest city was 36% above its fastest, against 89–219% through OpenRouter.
- City spread: 36% direct, 89–219% via OpenRouter
- Slow days: 0 direct, 1 via OpenRouter See the day-by-day chart →
- Finding 3 ### Most of the wait isn't the networkConnection setup was only 11–19 ms. Most latency came after reaching the provider.
- Network: 11–19 ms
- First token: 462–883 ms
[See where the time goes →](#time-split)
Time to first token, day by day #
A dot marks a day when every compared city was well above that provider’s usual level.
Show the numbers #
| Day | DeepSeek direct | OpenRouter · Cloudflare | OpenRouter · Together | OpenRouter · DigitalOcean |
|---|---|---|---|---|
| 27 August 2026 | — | — | 722 ms | 949 ms |
| 28 August 2026 | — | 646 ms | 661 ms | 498 ms |
| 29 August 2026 | — | — | 602 ms | 756 ms |
| 31 August 2026 | — | — | 580 ms | 650 ms |
| 1 September 2026 | — | — | 936 ms | 913 ms |
| 2 September 2026 | — | — | 1586 ms | 540 ms |
| 3 September 2026 | — | 796 ms | 784 ms | 827 ms |
| 4 September 2026 | — | 654 ms | — | 772 ms |
| 5 September 2026 | 714 ms | 533 ms | 765 ms | 639 ms |
| 6 September 2026 | 493 ms | 599 ms | — | 666 ms |
| 7 September 2026 | 617 ms | — | 875 ms | 799 ms |
| 8 September 2026 | 538 ms | 626 ms | 1057 ms | 805 ms |
| 9 September 2026 | 434 ms | 820 ms | 759 ms | 988 ms |
| 10 September 2026 | 781 ms | 743 ms | 829 ms | 764 ms |
| 11 September 2026 | 611 ms | 729 ms | 697 ms | 705 ms |
| 12 September 2026 | 645 ms | 873 ms | 859 ms | 736 ms |
| 13 September 2026 | 676 ms | 650 ms | 689 ms | 843 ms |
| 14 September 2026 | 802 ms | 748 ms | 757 ms | 792 ms |
| 15 September 2026 | 902 ms | 1274 ms | 687 ms | 752 ms |
| 16 September 2026 | 733 ms | 714 ms | — | 695 ms |
| 17 September 2026 | 594 ms | 745 ms | — | 861 ms |
| 18 September 2026 | 732 ms | 727 ms | 5566 ms ▲ | 806 ms |
| 19 September 2026 | 630 ms | 631 ms | 663 ms | 789 ms |
| 20 September 2026 | 619 ms | 725 ms | — | 845 ms |
| 21 September 2026 | 700 ms | 638 ms | 1138 ms | 722 ms |
| 22 September 2026 | 638 ms | 637 ms | 814 ms | 881 ms |
Does your location matter? #
DeepSeek direct was faster from Singapore and Mumbai in our tests.
- DeepSeek direct
- OpenRouter · Cloudflare
- OpenRouter · Together
- OpenRouter · DigitalOcean
Show exact values #
| City | DeepSeek direct | OpenRouter · Cloudflare | OpenRouter · Together | OpenRouter · DigitalOcean |
|---|---|---|---|---|
| Amsterdam | 733 msslower 1.1 s | 537 msslower 1.2 s | 722 msslower 3.3 s | 608 msslower 931 ms |
| San Francisco | 781 msslower 957 ms | 659 msslower 1.9 s | 627 msslower 5.8 s | 546 msslower 895 ms |
| Montreal | 732 msslower 989 ms | 643 msslower 1.5 s | 692 msslower 5.5 s | 422 msslower 2.1 s |
| Singapore | 621 msslower 866 ms | 782 msslower 1.8 s | 967 msslower 6.0 s | 939 msslower 2.1 s |
| Tokyo | 575 msslower 1.2 s | 784 msslower 2.1 s | 724 msslower 5.9 s | 1081 msslower 2.0 s |
| Mumbai | 664 msslower 883 ms | 1211 msslower 2.0 s | 1183 msslower 4.9 s | 1348 msslower 2.3 s |
Bold is the lowest typical time in that city. "Slower" is the time 95% of the daily requests from that city finished within.
Where does the wait happen? #
How much of the wait a closer server could remove, and how much is the provider’s own queue and model.
Direct vs OpenRouter #
9–22 September 2026 · six test locations
| Metric | Direct | Cloudflare | Together | DigitalOcean |
|---|---|---|---|---|
| Typical TTFT | 698 ms | 721 ms | 723 ms | 774 ms |
| Location spread | 36% | 126% | 89% | 219% |
Direct uses deepseek-v4-flash; the OpenRouter routes use deepseek-v4-flash-0731. Methodology →
Show full benchmark details ↓ #
| Detail | DeepSeek direct | OpenRouter · Cloudflare | OpenRouter · Together | OpenRouter · DigitalOcean |
|---|---|---|---|---|
| Model | deepseek-v4-flash | deepseek-v4-flash-0731 | deepseek-v4-flash-0731 | deepseek-v4-flash-0731 |
| Slower requests, p95 (median city) | 973 ms | 1820 ms | 5649 ms | 2039 ms | | Fastest location | Tokyo · 575 ms | Amsterdam · 537 ms | San Francisco · 627 ms | Montreal · 422 ms | | Slowest location | San Francisco · 781 ms | Mumbai · 1211 ms | Mumbai · 1183 ms | Mumbai · 1348 ms | | Days when every city slowed together | 0 | 0 | 1 | 0 | | Requests in this window | 420 over 14 daily runs | 417 over 14 daily runs | 326 over 11 daily runs | 419 over 14 daily runs | | Latest controlled run | 7 September 2026 | 20 September 2026 | 13 September 2026 | 20 September 2026 |
How we test #
We send the same streaming prompt from six cities to DeepSeek directly and through OpenRouter, then measure time to first visible token.
Same request on both sides: thinking off, temperature 0. The Fireworks and BaseTen routes are left out: one throttled us, the other never returned a visible token.
Good to know #
- Network setup is the hop to OpenRouter's edge (11–19 ms from every city); the hop from OpenRouter to the provider's machines is counted as waiting for the model, which is why Mumbai still waits longer than Montreal.
- Direct sends deepseek-v4-flash, which DeepSeek has served with V4.1 Flash since 10 September (V4 Flash is retired), while the OpenRouter routes pin deepseek-v4-flash-0731, so the two sides have not been the same model since that day.
- Each OpenRouter request pins one provider with fallbacks off, and a city's run is discarded if any response came from another provider.
- The prompt is "Say 'ok' and nothing else." (max_tokens 256, temperature 0, thinking off), sent daily between 03:30 and 06:46 UTC, 5 requests per city after one discarded warm-up.
- We time the first visible token only; tokens per second, long prompts and thinking on are not measured.
How does your API compare? #
Test your endpoint from the same 6 cities used here and see your response time next to these numbers.
Run a free latency test No account required · Takes about 30 seconds.