- OpenAI cut GPT-5.6 Luna to $0.20 per million input tokens and $1.20 per million output tokens on July 30, 2026. (openai.com)
- Anthropic made Claude Sonnet 5’s $2/$10 per-million-token price permanent on August 10, 2026; Opus 4.8 remains listed at $5/$25.
- DeepSeek’s V4 Flash lists a 1-million-token context window and prices of $0.14 per million cache-miss input tokens and $0.28 per million output tokens before its scheduled August 16 repricing. (api-docs.deepseek.com)
- Enterprise buyers can route different tasks to different models, but prompts, evaluations, guardrails and reliability testing create switching costs.
OpenAI and Anthropic are lowering prices for production-oriented models as Chinese competitors push API costs toward a level that makes frontier-model premiums harder to justify for routine workloads. The changes give enterprise developers cheaper options for coding, document processing and software agents, while widening the spread between U.S. providers’ premium offerings and Chinese alternatives. (openai.com)
OpenAI cut GPT-5.6 Luna’s API price by 80% on July 30, to $0.20 per million input tokens and $1.20 per million output tokens. It also cut GPT-5.6 Terra by 20%, to $2 per million input tokens and $12 per million output tokens. Both models remain available through the OpenAI API, ChatGPT Work and Codex. (openai.com)
OpenAI and Anthropic are compressing the middle of the market #
OpenAI’s flagship GPT-5.6 Sol remains priced at $5 per million input tokens and $30 per million output tokens. Before the July 30 reduction, the company listed Terra at $2.50/$15 and Luna at $1/$6; the new prices are $2/$12 for Terra and $0.20/$1.20 for Luna. (openai.com)
The three GPT-5.6 models have a 1.05-million-token context window and a maximum output of 128,000 tokens. OpenAI also introduced Fast mode for Sol, which it says can deliver up to 2.5 times the speed of standard processing at twice the price. Those are OpenAI’s published specifications and performance claims, not independent validation. (openai.com)
Anthropic has taken a similar position in the middle of its lineup. Sonnet 5 supports a 1-million-token context window and 128,000-token maximum output. Anthropic initially announced $2/$10 introductory pricing through August 31, followed by $3/$15 standard pricing, but changed course on August 10 and made the lower rate permanent.
Anthropic says Sonnet 5 improves on Sonnet 4.6 and can match Opus 4.8 on some tasks at higher effort levels. Its published list price for Opus 4.8 is $5 per million input tokens and $25 per million output tokens. The comparison is based on Anthropic’s evaluations and should not be read as a universal ranking across enterprise workloads.
DeepSeek establishes a much lower API benchmark #
DeepSeek’s V4 Flash and V4 Pro both support a 1-million-token context window and up to 384,000 output tokens. V4 Flash is listed at $0.14 per million cache-miss input tokens and $0.28 per million output tokens. V4 Pro is listed at $0.435 and $0.87, respectively; cache-hit input prices are lower. (api-docs.deepseek.com)
DeepSeek’s documentation says those rates will change at 16:00 UTC on August 16, 2026, when the company introduces peak and off-peak billing. During peak hours, V4 Flash will cost $0.44 per million input tokens and $1.32 per million output tokens, while V4 Pro will cost $1.32 and $3.96. The scheduled V4 Pro peak output price would still be below OpenAI Sol’s $30 and Anthropic Opus 4.8’s $25, although token prices do not capture differences in quality, latency, availability or support. (api-docs.deepseek.com)
Alibaba is also offering long-context models below the U.S. leaders in some regions. Alibaba Cloud lists Qwen3.7 Max internationally at $2.50 per million input tokens and $7.50 per million output tokens, with a 1-million-token context window. Its Chinese-mainland listing for Qwen3 Max starts at $0.359/$1.434 for requests up to 32,000 tokens and rises for longer inputs.
Chinese models are also attracting users outside China. The Associated Press reported that Mozilla CTO Raffi Krikorian moved many daily tasks to Moonshot’s Kimi K3 shortly after its launch. AP also reported that Bank of America analysts estimated Kimi K3’s price at roughly half that of OpenAI’s GPT-5.6 Sol, though that comparison concerns a different model and pricing structure than DeepSeek’s V4 Flash. (apnews.com)
Price is only one switching cost #
Lower API prices make model routing more attractive, but enterprise adoption is not determined by token cost alone. Teams must retest prompts, tool calls, output formats, safety filters, latency and failure rates before moving a production workflow. Andreessen Horowitz’s survey of enterprise CIOs found that companies were hesitant to switch after investing in guardrails, prompting and reliability work.
That cost is lower for workloads managed through a model gateway. Factory CEO Matan Grinberg told Axios that his company uses a router to select models by task and said buyers do not want to depend on one provider. Microsoft has also considered a hosted DeepSeek option for Copilot Cowork as part of a broader multi-model strategy.
The price cuts create a trade-off for providers: cheaper tokens may expand usage, but they also reduce revenue per unit unless efficiency gains offset the decline. OpenAI says GPT-5.6 kernel work reduced end-to-end serving costs by 20% and increased token-generation efficiency by more than 15%. Those figures are company claims; public disclosures do not establish the providers’ model-level margins.
The emerging enterprise pattern is therefore tiered rather than winner-take-all. A company may reserve Sol or Opus for difficult reasoning, use Sonnet 5 or Terra for general workflows, and route repetitive tasks to Luna, DeepSeek V4 Flash or Qwen. The commercial question is increasingly how much buyers will pay for reliability, safety, latency and completed-task quality—not simply which model has the lowest price per token.
Companies mentioned #
Further sources #
[[1] OpenAI, “Advancing the price-performance frontier with GPT-5.6,” July 30, 2026.… ↗](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
[[2] OpenAI, “GPT-5.6: Frontier intelligence that scales with your ambition,” July 9… ↗](https://openai.com/index/gpt-5-6/)
[3] OpenAI API documentation, “Compare models.” Context windows, output limits and … ↗
[[4] Anthropic, “Introducing Claude Sonnet 5,” June 30, 2026, updated August 10. Son… ↗](https://www.anthropic.com/news/claude-sonnet-5)
[[5] Anthropic Claude Platform documentation, “What’s new in Claude Sonnet 5.” Conte… ↗](https://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5)
[[6] DeepSeek API documentation, “Models & Pricing.” V4 model specifications, curren… ↗](https://api-docs.deepseek.com/quick_start/pricing/)+6 more
The stories that matter, in one email. Free — unsubscribe anytime.