{"slug": "deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization", "title": "DeepSeek V4 Flash at 278 tok/s, full precision, no quantization", "summary": "RunInfra lists DeepSeek V4 Flash, an LLM served as deepseek-ai/DeepSeek-V4-Flash-0731, at $0.13 per 1M input tokens and $0.27 per 1M output tokens, with a 1,048,576-token context window and OpenAI-compatible chat completions. The model achieves 278 tokens per second at full precision with no quantization, according to the headline.", "body_md": "`deepseek-ai/DeepSeek-V4-Flash-0731`\n\nDeepSeek V4 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as deepseek-ai/DeepSeek-V4-Flash-0731 at $0.13 per 1M input tokens and $0.27 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions.\n\nUSD, pay per token\n\nHosted inference is available to every account. You pay for token usage from the same workspace balance.\n\nConfirm how your client reaches this model.\n\nCheck the limits your workload must fit.\n\nSet RUNINFRA_GATEWAY_KEY to your workspace API key before using an example.\n\nVerify the company and operating credentials behind this API.\n\n© 2026 RunInfra. All rights reserved.\n\nSee which request modes the API supports.\n\nCheck what the endpoint adds when a request omits a field.", "url": "https://wpnews.pro/news/deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization", "canonical_source": "https://runinfra.ai/inference-api/deepseek-v4-flash", "published_at": "2026-08-15 13:26:44+00:00", "updated_at": "2026-08-15 13:40:53.054757+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-infrastructure"], "entities": ["RunInfra", "DeepSeek", "DeepSeek-V4-Flash-0731"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization", "markdown": "https://wpnews.pro/news/deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization.md", "text": "https://wpnews.pro/news/deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-flash-at-278-tok-s-full-precision-no-quantization.jsonld"}}