cd /news/artificial-intelligence/deepseek-v4-pricing-now-depends-on-w… · home topics artificial-intelligence article
[ARTICLE · art-98646] src=dev.to ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DeepSeek V4 pricing now depends on what time you run it

DeepSeek moved its entire API to peak and off-peak billing on August 16, 2026, with flagship V4-Pro output costing $3.96 per million tokens at peak and $1.98 off-peak, up from a flat $0.87. The schedule, which includes windows from 01:00-04:00 and 06:00-10:00 UTC, means European teams face peak pricing during their working hours while American teams pay off-peak rates throughout the business day. Cached input prices rose by as much as 1,100% at peak, disproportionately affecting applications that resend large system prompts.

read5 min views13 publishedAug 16, 2026

DeepSeek moved its entire API over to peak and off-peak billing at 16:00 UTC on August 16, 2026. Output generated by the flagship V4-Pro model now costs $3.96 per million tokens at peak, and $1.98 the rest of the day, per DeepSeek's own pricing page. The previous promotional rate was a flat $0.87 at any hour of the day.

Every outlet reported the multiple. Almost nobody reported the schedule, and the schedule is what determines your invoice. Peak hours are 01:00-04:00 and 06:00-10:00 UTC. That is seven hours out of every 24. Whether those expensive hours overlap your working day or your sleep depends entirely on your time zone, and the answer separates American teams from European ones.

DeepSeek is a Chinese artificial intelligence lab that sells access to its own large language models through an OpenAI-compatible API. Two models are generally available. V4-Flash is the cheaper option. V4-Pro-0813 is the flagship, and it arrived on August 13, 2026.

Peak pricing is a billing arrangement that charges double for identical work performed during specified hours. Both models adopted it simultaneously at the August cutover. The rates below are in US dollars per million tokens, and they come from the DeepSeek pricing table for both models.

Model and rate Cache miss input Cache hit input Output
V4-Pro, old flat rate $0.435 $0.003625 $0.87
V4-Pro, off-peak $0.66 $0.022 $1.98
V4-Pro, peak $1.32 $0.044 $3.96
V4-Flash, old flat rate $0.14 $0.0028 $0.28
V4-Flash, off-peak $0.22 $0.007 $0.66
V4-Flash, peak $0.44 $0.014 $1.32

DeepSeek attributes the restructuring to infrastructure capacity rather than margin. "To allocate resources more reasonably, we will adopt peak/off-peak pricing," the firm said in the notice quoted by PYMNTS, "encouraging users to schedule their tasks based on actual usage."

Examine the cache hit column carefully, because it contains the largest multiplier in the entire announcement. Cached input is the discounted rate applied to prompt prefixes the provider has already processed and stored. On V4-Pro that rate climbed from $0.003625 to $0.044 at peak. That is a rise of about 1,100%, as InfoWorld set out in its rate card breakdown. Applications that resend a large fixed system prompt on every request absorb that increase disproportionately.

Peak billing applies during 01:00-04:00 and 06:00-10:00 UTC, and everything outside those windows bills at half price. Converted into local time for August 2026, with daylight saving observed on both sides of the Atlantic, the windows are distributed as follows.

| Region | First peak window | Second peak window |

|---|---|---|
| US Eastern (EDT, UTC-4) | 21:00-00:00 | 02:00-06:00 |
| US Central (CDT, UTC-5) | 20:00-23:00 | 01:00-05:00 |
| US Pacific (PDT, UTC-7) | 18:00-21:00 | 23:00-03:00 |
| United Kingdom (BST, UTC+1) | 02:00-05:00 | 07:00-11:00 |
| Central Europe (CEST, UTC+2) | 03:00-06:00 | 08:00-12:00 |

Consider the European rows first. In Berlin, Paris, Madrid and Amsterdam the second window covers 08:00 until 12:00, which is exactly the working morning. In London it covers 07:00 until 11:00, capturing stand-up, the morning deployment and most interactive development. Nobody in those offices selected that arrangement.

The American rows invert the outcome. Nothing between 06:00 and 21:00 Eastern falls inside a peak window, on the same UTC schedule DeepSeek published. A team in New York, Chicago or Austin therefore pays the discounted rate throughout its business day, with no additional cron entry and no code change.

Consider a deliberately modest workload, such as 20 million output tokens a month on V4-Pro, the model Quartz reported as generally available across app, web and API on August 13, 2026. That volume represents a few thousand substantial generations, not a hyperscale deployment.

American teams arrive at the middle number accidentally. European teams arrive at the highest number, equally accidentally. The gap between $39.60 and $79.20 is the practical consequence of this announcement, and it becomes visible only when the UTC schedule is converted into local working hours.

One old rule still holds, and this blog has made the case before: the cheapest AI API is not the cheapest to run. A published per-token rate describes what an individual token costs. Your invoice reflects tokens consumed per task, retry behavior, cached prefix size, and now the position of the clock when your scheduler fires.

DeepSeek remains dramatically cheaper than the American laboratories, even after quadrupling its rates. These are list prices per million tokens, as set out in coverage of the Kimi K3 launch.

Model Output price Input price
DeepSeek V4-Pro, off-peak $1.98 $0.66
DeepSeek V4-Pro, peak $3.96 $1.32
Moonshot Kimi K3 $15 $3
OpenAI GPT-5.6 Sol $30 $5
Anthropic Fable 5 $50 $10

At its most expensive hour, V4-Pro output remains roughly one eighth the price of GPT-5.6 Sol, and approximately one thirteenth the price of Fable 5. The competitive tier has therefore not changed. What changed is the scheduling arithmetic inside that tier, alongside the disappearance of a promotional rate that made the model appear almost free.

One caveat deserves stating explicitly. Price per token is not price per task. A model that reasons verbosely can emit three times the output tokens of a concise competitor while completing identical work. Benchmark finished tasks against your own workload before migrating production traffic.

The appropriate response depends on where your requests originate and when they execute.

That final scenario connects to a question worth resolving beforehand: does the free AI API tier train on your data? Pricing is one dimension of provider selection. Data retention is the other dimension, and a retention decision is considerably harder to reverse than a billing decision.

Sources: DeepSeek API pricing docs, InfoWorld on the size of the rise, PYMNTS on the peak-hour rule, Quartz on the V4-Pro launch, The Decoder on rival model prices.

Originally published on www.nihardaily.com. For more articles like this one, visit www.nihardaily.com.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepseek-v4-pricing-…] indexed:0 read:5min 2026-08-16 ·