DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization RunInfra lists DeepSeek V4 Pro as deepseek-ai/DeepSeek-V4-Pro-0813, an LLM with a 1,048,576-token context window and a throughput of 207 tokens per second with no quantization. The API, priced at $0.60 per 1M input tokens and $1.90 per 1M output tokens, offers OpenAI-compatible chat completions. deepseek-ai/DeepSeek-V4-Pro-0813 DeepSeek V4 Pro is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Pro-0813 at $0.60 per 1M input tokens and $1.90 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions. USD, pay per token Confirm how your client reaches this model. Check the limits your workload must fit. See which request modes the API supports. Set RUNINFRA GATEWAY KEY to your workspace API key before using an example. Verify the company and operating credentials behind this API. © 2026 RunInfra. All rights reserved.