00:00
2026-08-14
oskrim.github.io
large-language-models
Deepseek V4 Flash 0731 latency numbers from nine providers
A one-time snapshot of DeepSeek V4 Flash 0731 latency across nine inference providers found Baseten fastest with 3,980 decode tok/s and 7.72 s total p99, while Azure ran an older checkpoint and Scalewβ¦