# DeepSeek-V4-Flash 2.98x faster on 4x B200, lossless

> Source: <https://twitter.com/Akashi203/status/2084373935454400964>
> Published: 2026-08-03 20:34:55+00:00

We optimized DeepSeek-V4-Flash-0731 at @runinfrai and made it 2.98x faster than baseline. Median latency dropped from 9248 ms to 3095 ms and throughput went from 113 tokens/s to 363 tokens/s. The optimization is lossless, with output parity by construction
We don't host it. You run it on your own hardware.
It costs $60 one time and it is yours forever
The kit is tuned for vLLM 0.25.0 on 4x B200 and ships with the full benchmark receipt. Deploy it on any GPU provider or on your own machines
No deprecations, no silent model swaps, no per-token pricing. You control your intelligence!!
runinfra.ai/catalog/deepse…
