{"slug": "deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization", "title": "DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization", "summary": "RunInfra lists DeepSeek V4 Pro as deepseek-ai/DeepSeek-V4-Pro-0813, an LLM with a 1,048,576-token context window and a throughput of 207 tokens per second with no quantization. The API, priced at $0.60 per 1M input tokens and $1.90 per 1M output tokens, offers OpenAI-compatible chat completions.", "body_md": "`deepseek-ai/DeepSeek-V4-Pro-0813`\n\nDeepSeek V4 Pro is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Pro-0813 at $0.60 per 1M input tokens and $1.90 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions.\n\nUSD, pay per token\n\nConfirm how your client reaches this model.\n\nCheck the limits your workload must fit.\n\nSee which request modes the API supports.\n\nSet RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.\n\nVerify the company and operating credentials behind this API.\n\n© 2026 RunInfra. All rights reserved.", "url": "https://wpnews.pro/news/deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization", "canonical_source": "https://runinfra.ai/inference-api/deepseek-v4-pro", "published_at": "2026-08-17 01:55:09+00:00", "updated_at": "2026-08-17 02:10:39.219052+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-infrastructure"], "entities": ["RunInfra", "DeepSeek", "DeepSeek-V4-Pro-0813"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization", "markdown": "https://wpnews.pro/news/deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization.md", "text": "https://wpnews.pro/news/deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-pro-at-207-tok-s-with-the-full-1m-context-no-quantization.jsonld"}}