DeepSeek Warns Its Ultra-Cheap V4-Flash Pricing Won't Last DeepSeek's official API documentation lists V4-Flash at $0.14 per million cache-miss input tokens, $0.28 per million output tokens, and $0.0028 per million cache-hit input tokens, making it unusually cheap for a 1 million-token context model, but the company's pricing page reserves the right to adjust prices without specifying a date. The model, released in April 2026 with 284 billion total parameters and 13 billion active parameters, scores 40 on Artificial Analysis's Intelligence Index and generates output at 118.2 tokens per second, with weights available on Hugging Face under an MIT license. DeepSeek's April 24 change log indicates that old deepseek-chat and deepseek-reasoner names will map to V4-Flash compatibility modes until July 24, 2026, and developers should model potential price changes in production costs. DeepSeek V4-Flash is still unusually cheap, but the published facts don't support the claim that a named founder warned developers about an imminent price hike. DeepSeek's real pricing story is simpler, and more useful for you. The company's official API docs list V4-Flash at $0.14 per million cache-miss input tokens, $0.28 per million output tokens, and $0.0028 per million cache-hit input tokens. That is extremely low for a 1 million-token context model with tool calls, JSON output, and both thinking and non-thinking modes. That is the story. The part that needed cleaning up was the drama around it. DeepSeek's own pricing page says product prices may vary and that the company reserves the right to adjust them. It doesn't give a date for a new V4-Flash price. It doesn't publish a fresh August warning saying the current rate is about to disappear. It also doesn't support the claim that a founder named Jun Song defended a 2x to 10x increase on X. DeepSeek's widely reported founder is Liang Wenfeng, and that specific attribution doesn't check out. You should still pay attention to the price. According to DeepSeek's April 24 change log, the API began supporting DeepSeek-V4-Pro and DeepSeek-V4-Flash through both OpenAI-compatible and Anthropic-compatible interfaces. The same notice gave developers three months to move away from the old deepseek-chat and deepseek-reasoner names. Those names still work, for now: DeepSeek says they'll map to V4-Flash compatibility modes until July 24, 2026. The cheap model is real Artificial Analysis lists DeepSeek V4-Flash, in reasoning mode at max effort, as an April 2026 release with 284 billion total parameters and 13 billion active parameters. The model scores 40 on the firm's Intelligence Index, above the comparable open-weight median shown on the same page. It also generates output at 118.2 tokens per second on DeepSeek's API, which is fast enough that the low token price isn't hiding a uselessly slow service. No mystery there. DeepSeek is using price as a weapon. For a developer, the practical point is not whether V4-Flash beats every closed model on every benchmark. It doesn't need to. A model that is good enough for routing, drafting, coding help, classification, research cleanup, and long-context document work becomes hard to ignore when the input price starts at fourteen cents per million tokens. Cache hits at $0.0028 make the gap sharper if your workload repeats system prompts, templates, policies, or large reference blocks. Use it if it fits. The open-weight detail matters too. Artificial Analysis lists the model weights as available on Hugging Face under an MIT license. That gives DeepSeek a different kind of pressure point from an API-only provider. You can use the hosted API when it is cheap and convenient, then at least evaluate self-hosting or third-party routing if your usage grows large enough to justify the work. The risk sits in the invoice DeepSeek's official docs keep one plain warning in view: prices may change. You shouldn't turn that into a fake countdown, but you also shouldn't ignore it. Any team building production costs around today's V4-Flash rates needs to model a higher bill, because the provider has given itself room to move. That is the real catch. The April model shift also shows how quickly DeepSeek is willing to move developers from one naming scheme to another. The old chat and reasoner endpoints were not treated as permanent products. They were folded into V4-Flash modes, with a dated deprecation notice and a path for migration. If your app hard-codes model names, prices, and behavior assumptions, DeepSeek has already shown you why that is fragile. Frankly, the right answer is boring. Test V4-Flash on your real prompts, not on someone else's leaderboard. Measure cache hit rates. Compare total task cost, including retries and long outputs. Then build a pricing buffer before you point serious production traffic at it. Price is the weapon. But a cheap AI model only saves money when it finishes the job well enough that you don't have to run the same task twice. DeepSeek has not proved that the current V4-Flash price is permanent. It has proved something more immediate: at the published API rates, the model is cheap enough that developers have to benchmark it against their own stack. If the numbers hold, you're ahead. If they move, you want to know before the invoice does. Also read: An AI Assistant Booking a Gym Class in Melbourne Ended Up Hacking the Site https://startupfortune.com/an-ai-assistant-booking-a-gym-class-in-melbourne-ended-up-hacking-the-site/ • Astera Labs Guided to $550 Million and Wall Street Sold the Stock Anyway https://startupfortune.com/astera-labs-guided-to-550-million-and-wall-street-sold-the-stock-anyway/ • Situational Awareness Pours $400 Million Into Chip Startup Source Foundry https://startupfortune.com/situational-awareness-pours-400-million-into-chip-startup-source-foundry/