{"slug": "deepseek-v4-pricing-now-depends-on-what-time-you-run-it", "title": "DeepSeek V4 pricing now depends on what time you run it", "summary": "DeepSeek moved its entire API to peak and off-peak billing on August 16, 2026, with flagship V4-Pro output costing $3.96 per million tokens at peak and $1.98 off-peak, up from a flat $0.87. The schedule, which includes windows from 01:00-04:00 and 06:00-10:00 UTC, means European teams face peak pricing during their working hours while American teams pay off-peak rates throughout the business day. Cached input prices rose by as much as 1,100% at peak, disproportionately affecting applications that resend large system prompts.", "body_md": "DeepSeek moved its entire API over to peak and off-peak billing at 16:00 UTC on August 16, 2026. Output generated by the flagship V4-Pro model now costs $3.96 per million tokens at peak, and $1.98 the rest of the day, per [DeepSeek's own pricing page](https://api-docs.deepseek.com/quick_start/pricing). The previous promotional rate was a flat $0.87 at any hour of the day.\n\nEvery outlet reported the multiple. Almost nobody reported the schedule, and the schedule is what determines your invoice. Peak hours are 01:00-04:00 and 06:00-10:00 UTC. That is seven hours out of every 24. Whether those expensive hours overlap your working day or your sleep depends entirely on your time zone, and the answer separates American teams from European ones.\n\n**DeepSeek** is a Chinese artificial intelligence lab that sells access to its own large language models through an OpenAI-compatible API. Two models are generally available. V4-Flash is the cheaper option. V4-Pro-0813 is the flagship, and it arrived on August 13, 2026.\n\n**Peak pricing** is a billing arrangement that charges double for identical work performed during specified hours. Both models adopted it simultaneously at the August cutover. The rates below are in US dollars per million tokens, and they come from the [DeepSeek pricing table for both models](https://api-docs.deepseek.com/quick_start/pricing).\n\n| Model and rate | Cache miss input | Cache hit input | Output |\n|---|---|---|---|\n| V4-Pro, old flat rate | $0.435 | $0.003625 | $0.87 |\n| V4-Pro, off-peak | $0.66 | $0.022 | $1.98 |\n| V4-Pro, peak | $1.32 | $0.044 | $3.96 |\n| V4-Flash, old flat rate | $0.14 | $0.0028 | $0.28 |\n| V4-Flash, off-peak | $0.22 | $0.007 | $0.66 |\n| V4-Flash, peak | $0.44 | $0.014 | $1.32 |\n\nDeepSeek attributes the restructuring to infrastructure capacity rather than margin. \"To allocate resources more reasonably, we will adopt peak/off-peak pricing,\" the firm said in the notice [quoted by PYMNTS](https://www.pymnts.com/news/artificial-intelligence/2026/deepseek-introduces-peak-hour-pricing-that-quadruples-current-levels/), \"encouraging users to schedule their tasks based on actual usage.\"\n\nExamine the cache hit column carefully, because it contains the largest multiplier in the entire announcement. **Cached input** is the discounted rate applied to prompt prefixes the provider has already processed and stored. On V4-Pro that rate climbed from $0.003625 to $0.044 at peak. That is a rise of about 1,100%, as [InfoWorld set out in its rate card breakdown](https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html). Applications that resend a large fixed system prompt on every request absorb that increase disproportionately.\n\nPeak billing applies during 01:00-04:00 and 06:00-10:00 UTC, and everything outside those windows bills at half price. Converted into local time for August 2026, with daylight saving observed on both sides of the Atlantic, the windows are distributed as follows.\n\n| Region | First peak window | Second peak window |\n|---|---|---|\n| US Eastern (EDT, UTC-4) | 21:00-00:00 | 02:00-06:00 |\n| US Central (CDT, UTC-5) | 20:00-23:00 | 01:00-05:00 |\n| US Pacific (PDT, UTC-7) | 18:00-21:00 | 23:00-03:00 |\n| United Kingdom (BST, UTC+1) | 02:00-05:00 | 07:00-11:00 |\n| Central Europe (CEST, UTC+2) | 03:00-06:00 | 08:00-12:00 |\n\nConsider the European rows first. In Berlin, Paris, Madrid and Amsterdam the second window covers 08:00 until 12:00, which is exactly the working morning. In London it covers 07:00 until 11:00, capturing stand-up, the morning deployment and most interactive development. Nobody in those offices selected that arrangement.\n\nThe American rows invert the outcome. Nothing between 06:00 and 21:00 Eastern falls inside a peak window, on the same UTC schedule DeepSeek published. A team in New York, Chicago or Austin therefore pays the discounted rate throughout its business day, with no additional cron entry and no code change.\n\nConsider a deliberately modest workload, such as 20 million output tokens a month on V4-Pro, the model [Quartz reported as generally available across app, web and API](https://qz.com/deepseek-v4-pro-official-launch-081326) on August 13, 2026. That volume represents a few thousand substantial generations, not a hyperscale deployment.\n\nAmerican teams arrive at the middle number accidentally. European teams arrive at the highest number, equally accidentally. The gap between $39.60 and $79.20 is the practical consequence of this announcement, and it becomes visible only when the UTC schedule is converted into local working hours.\n\nOne old rule still holds, and this blog has made the case before: [the cheapest AI API is not the cheapest to run](https://www.nihardaily.com/posts/the-cheapest-ai-api-is-not-the-cheapest-to-run). A published per-token rate describes what an individual token costs. Your invoice reflects tokens consumed per task, retry behavior, cached prefix size, and now the position of the clock when your scheduler fires.\n\nDeepSeek remains dramatically cheaper than the American laboratories, even after quadrupling its rates. These are list prices per million tokens, [as set out in coverage of the Kimi K3 launch](https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/).\n\n| Model | Output price | Input price |\n|---|---|---|\n| DeepSeek V4-Pro, off-peak | $1.98 | $0.66 |\n| DeepSeek V4-Pro, peak | $3.96 | $1.32 |\n| Moonshot Kimi K3 | $15 | $3 |\n| OpenAI GPT-5.6 Sol | $30 | $5 |\n| Anthropic Fable 5 | $50 | $10 |\n\nAt its most expensive hour, V4-Pro output remains roughly one eighth the price of GPT-5.6 Sol, and approximately one thirteenth the price of Fable 5. The competitive tier has therefore not changed. What changed is the scheduling arithmetic inside that tier, alongside the disappearance of a promotional rate that made the model appear almost free.\n\nOne caveat deserves stating explicitly. Price per token is not price per task. A model that reasons verbosely can emit three times the output tokens of a concise competitor while completing identical work. Benchmark finished tasks against your own workload before migrating production traffic.\n\nThe appropriate response depends on where your requests originate and when they execute.\n\nThat final scenario connects to a question worth resolving beforehand: [does the free AI API tier train on your data?](https://www.nihardaily.com/posts/does-the-free-ai-api-tier-train-on-your-data) Pricing is one dimension of provider selection. Data retention is the other dimension, and a retention decision is considerably harder to reverse than a billing decision.\n\n**Sources:** [DeepSeek API pricing docs](https://api-docs.deepseek.com/quick_start/pricing), [InfoWorld on the size of the rise](https://www.infoworld.com/article/4209439/deepseek-raises-some-v4-prices-by-more-than-10x-as-ai-demand-strains-capacity.html), [PYMNTS on the peak-hour rule](https://www.pymnts.com/news/artificial-intelligence/2026/deepseek-introduces-peak-hour-pricing-that-quadruples-current-levels/), [Quartz on the V4-Pro launch](https://qz.com/deepseek-v4-pro-official-launch-081326), [The Decoder on rival model prices](https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/).\n\n*Originally published on www.nihardaily.com. For more articles like this one, visit www.nihardaily.com.*", "url": "https://wpnews.pro/news/deepseek-v4-pricing-now-depends-on-what-time-you-run-it", "canonical_source": "https://dev.to/akashdas/deepseek-v4-pricing-now-depends-on-what-time-you-run-it-1081", "published_at": "2026-08-16 09:53:13+00:00", "updated_at": "2026-08-16 10:42:03.971555+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-infrastructure"], "entities": ["DeepSeek", "V4-Pro", "V4-Flash", "PYMNTS", "InfoWorld"], "alternates": {"html": "https://wpnews.pro/news/deepseek-v4-pricing-now-depends-on-what-time-you-run-it", "markdown": "https://wpnews.pro/news/deepseek-v4-pricing-now-depends-on-what-time-you-run-it.md", "text": "https://wpnews.pro/news/deepseek-v4-pricing-now-depends-on-what-time-you-run-it.txt", "jsonld": "https://wpnews.pro/news/deepseek-v4-pricing-now-depends-on-what-time-you-run-it.jsonld"}}