DeepSeek V4.1 Flash: Pro Routing, Prices and Early Tests DeepSeek announced via a September 9 customer email that it will release V4.1 Flash around September 10, Beijing time, temporarily routing Pro requests to Flash and introducing new USD pricing per million tokens effective September 10 at 04:00 UTC, with cached input at $0.003 off-peak and $0.006 peak, uncached input at $0.15 off-peak and $0.30 peak, and output at $0.60 off-peak and $1.20 peak. The notice does not provide a V4.1 Pro launch date or a final callable model identifier, and DeepSeek's public documentation still maps the Pro alias to V4-Pro-0813 as of September 9. DeepSeek V4.1 Flash is an announced change to the model behind Pro requests, as well as a cheaper option to evaluate. A September 9 customer email supplied to Digital Applied says DeepSeek plans to release it around September 10, Beijing time, and temporarily route Pro requests to Flash. For a team already using Pro, the immediate job is to save a comparison of its current results before that transition. There is an early public design evaluation worth examining, but we have not found a V4.1 technical report or official numerical benchmark suite in DeepSeek’s public release documentation. This article separates the customer notice, the published evaluation and our recommendations. It does not establish that the final model has launched or that it improves every workload. 1. 01Prepare for a backend change.The notice describes temporary Pro-to-Flash routing after launch, pending V4.1 Pro. An unchanged model selection would not establish unchanged behavior. 2. 02Keep the two deadlines separate.The release timing is approximate. The announced billing change is September 10 at 04:00 UTC. 3. 03Test the work you actually run.The public design evaluation offers a useful starting point. Its overall ranking does not settle performance on every task. 01 — Announced billingThe prices in the customer notice The email gives the following USD rates per million tokens, effective September 10 at 04:00 UTC, or noon in Beijing. Cached input means input billed at the cache-hit rate; uncached input uses the cache-miss rate. These are announced future prices, not a claim that the public price list has already changed. | Announced DeepSeek Flash USD prices per million tokens, effective September 10, 2026 at 04:00 UTC, from the September 9 customer notice. | | | |---|---|---| | Token category | Off-peak | Peak | |---|---|---| | Cached input | $0.003 | $0.006 | | Uncached input | $0.15 | $0.30 | | Output | $0.60 | $1.20 | Peak hours are Monday through Friday, 01:00–04:00 and 06:00–10:00 UTC. All other hours are off-peak. The existing DeepSeek pricing documentation https://api-docs.deepseek.com/quick start/pricing confirms those windows, but still displayed the previous rates at our September 9 check. Do not substitute local clock time for UTC when estimating a scheduled job. A lower token rate is only one part of a lower bill. A replacement that takes more attempts, writes longer answers or needs additional review can erase some savings. For the broader scheduling issue, our off-peak LLM pricing analysis /blog/off-peak-llm-pricing-deepseek-glm-windows-2026 explains why the cheaper window and the cheaper completed task are different comparisons. 02 — Existing customersA Pro selection may serve a different model According to the notice, after V4.1 Flash launches and before V4.1 Pro arrives, Pro requests will be served by V4.1 Flash and charged at Flash rates. The notice supplies no V4.1 Pro launch date. It also does not establish the final callable V4.1 identifier, so there is no new configuration string to copy from this article. DeepSeek’s current API introduction https://api-docs.deepseek.com/ still maps the Pro alias to V4-Pro-0813. That is the existing documented arrangement; the email describes a forthcoming one. Our coverage of V4 Pro’s earlier general release /blog/deepseek-v4-pro-ga-official-release-2026 provides the dated background. The practical consequence is reproducibility. A saved prompt and the same model label may no longer reproduce the same underlying conditions. Record the request date, requested model and returned version information where available. If retaining the old model is essential, seek explicit confirmation of that option; neither the notice nor the public documentation we inspected establishes one. 03 — Early evidenceWhat the design benchmark can tell us OpenDesign Arena https://open-design.ai/llm-arena-for-design/ publishes these results for prototype-design tasks. They are the evaluator’s reported measurements, not tests run by Digital Applied or confirmation of the final production checkpoint. | Selected OpenDesign Arena results read September 9, 2026. Costs are estimated per artifact, not invoices. | | | | |---|---|---|---| | Model label | Mean score /100 | Mean minutes | Estimated USD/artifact | |---|---|---|---| | DeepSeek V4.1 Flash | 81.2 | 5.3 | $0.023 | | DeepSeek V4 Pro | 72.9 | 17.7 | $0.061 | | GPT-6 Astra | 82.7 | 11.1 | $1.61 | The dashboard subset reverses the overall DeepSeek ranking: V4.1 Flash scores 76.1 against Pro’s 83.0. A favorable average therefore does not mean every scenario improves. - Scoring - Five prototype scenarios; 30 points for meeting requirements and 70 for design quality. Non-rendering artifacts receive zero. - Costs and time - Costs use recorded token usage and list prices; they are not invoices. Timing excludes queue time. Do not equate these estimates with tomorrow’s announced tariff. - Limits - The inspected method does not establish exact checkpoint IDs, per-model sample counts or uncertainty intervals. These results concern prototypes, not general capability or production readiness. DeepSeek’s customer notice makes a broader superiority claim, including performance, speed and task completion time. We have not found the accompanying numerical suite in its public changelog https://api-docs.deepseek.com/updates/ . Older V4 scores cannot fill that gap: they describe different releases. Treat the notice as the company’s claim and the design evaluation as one reason to investigate it. 04 — Practical preparationSave a baseline before the transition Our recommendation is a small comparison using work you can judge. Choose examples with a known acceptance condition: an existing test suite, a required output format, a document answer checked against its source, or a prototype with specified interactions. Include a troublesome example alongside routine work. Attractive output alone is a weak acceptance test. 1. Save the current conditions. Keep prompts, input files, tool versions, permissions and model settings with the results. Record failures as well as successes. 2. Repeat after the change is confirmed. Keep the surrounding workflow consistent. If you change tools or instructions too, report that as a different experiment. 3. Count accepted outcomes. Record elapsed time, billed usage, retries and necessary human corrections. Compare total spend with the number of results you would actually use. For example, a generated dashboard can render successfully while showing the wrong totals or dropping a filter state. Write those checks before looking at the new output. This makes the comparison useful even when the replacement produces a more polished first impression. Our Astra and Fable comparison /blog/gpt-6-astra-vs-claude-fable-5-1-comparison develops the same distinction between advertised prices and the work needed to reach an accepted result. Evidence checked September 9, 2026: the supplied customer email, DeepSeek’s public API documentation and OpenDesign’s evaluation. The routing notice was not independently retrieved from the sign-in-only platform. Final V4.1 specifications, weights and general benchmark results remain unverified here. 05 — Today’s decisionPrepare the comparison, then judge the result Treat the replacement as a model change worth testing. Existing Pro users have a reason to capture today’s behavior and check the announced billing transition. New users have a promising candidate to evaluate once launch details are confirmed. Neither group needs to assume that a lower price or a higher aggregate score settles the quality question for its own work. Our AI transformation team /services/ai-transformation helps define acceptance tests and compare the cost of completed workflows before a wider rollout.