# OpenAI slashes GPT-5.6 Luna prices by 80% as Chinese rivals rewrite the AI cost equation

> Source: <https://startupfortune.com/openai-slashes-gpt-56-luna-prices-by-80-as-chinese-rivals-rewrite-the-ai-cost-equation/>
> Published: 2026-07-31 13:48:12+00:00

*OpenAI's Luna price cut is current, but the original article leaned on several unsupported rankings and market-share claims. The cleaner story is simpler: cheaper Chinese models forced OpenAI to defend the middle of its product line.*

Three weeks. That's how long it took OpenAI to cut the price of GPT-5.6 Luna after launching the GPT-5.6 family on July 9. Axios reported that the company reduced Luna to $0.20 per million input tokens and $1.20 per million output tokens on July 30, down from $1 and $6. Terra also came down, while Sol, the flagship model, stayed at $5 per million input tokens and $30 per million output tokens.

That is not generosity. Price is strategy. OpenAI can still charge frontier rates for Sol because customers will pay for the hardest coding, science, cybersecurity and agentic work when the model is clearly ahead. Luna is different. It sits in the part of the market where you can lose a customer one routing decision at a time.

OpenAI's own July 9 launch page described Luna as its most cost-efficient GPT-5.6 model, built for high-volume workloads. That phrasing matters because high volume is exactly where AI bills get ugly. A chatbot, coding assistant or internal agent doesn't just call a model once - it loops, retries, summarizes, calls tools. Tokens burn fast. And on a per-million-token price card, the whole thing can look deceptively cheap right up until it isn't.

## The Cheap Model Is Now The Benchmark

The competitive pressure is real. The Associated Press reported this week that U.S. developers and companies are increasingly testing Chinese models such as Moonshot's Kimi K3, Z.ai's GLM-5.2 and DeepSeek because they are cheaper, open-weight, and good enough for many production tasks. That last phrase is doing the work. Good enough at one-tenth the price changes buying behavior faster than another benchmark record.

You don't need to believe every leaderboard to see the direction of travel. Artificial Analysis currently lists GLM-5.2 with an Intelligence Index score of 51 and MiniMax-M3 at 44, with both released in June 2026. Its page for GPT-5.6 Luna, still showing the earlier $1 and $6 pricing when checked, puts Luna's score at 46. The exact rank will move as pricing pages update. The underlying point won't. OpenAI has now made Luna cheap enough that finance teams have to rerun the comparison.

DeepSeek has already shown how quickly share can move when the cost side is compelling. OpenRouter's own June blog said DeepSeek had climbed to nearly 20% of token share by early June and had been the top model author on its platform since mid-May. That is a practical signal, not a press-release signal. Developers route work where it runs cheaply and reliably.

Frankly, OpenAI had to answer that. If you are building with AI at scale, you don't care which lab has the prettier launch page when the monthly bill lands. You care whether the model is strong enough for the task and cheap enough to use without rationing every workflow.

## Budgets Are Forcing The Issue

Uber is the example to watch because it makes the cost problem concrete. TechCrunch reported in June, citing Bloomberg and The Information, that Uber introduced a $1,500 monthly cap per employee and per agentic coding tool after burning through its annual AI budget in four months. The tools included Anthropic's Claude Code and Cursor. That is what adoption looks like when it escapes the pilot program.

The point isn't that Uber made a mistake. The point is that employees used the tools heavily because the tools were useful. Once agentic coding becomes normal, token consumption stops behaving like a neat software subscription. One engineer asking for a small completion and another running several agents across a large codebase are not consuming the same product in any meaningful financial sense.

That is where Luna's cut bites. At $1 per million input tokens, OpenAI had a lower-cost option but not a shockingly cheap one. At $0.20, it can meet cost-sensitive workloads much closer to where Chinese open-weight models have been applying pressure. You still have to compare output prices, latency, context window, privacy considerations and vendor lock-in. You should. But the old assumption that OpenAI's serious models automatically sit outside the low-cost conversation is weaker now.

Sol's unchanged price is the other half of the message. OpenAI is not cutting everything because it doesn't think everything is commoditized. It is protecting the frontier while conceding the middle. That is a more honest posture than pretending one model family can hold every customer at every price point.

The next move belongs to buyers as much as vendors. If you run a startup or an enterprise AI team, the useful question is no longer which model is best in the abstract. Ask which tasks deserve Sol, which can run on Luna, and which should be routed to a Chinese open-weight model or another cheaper system. The companies that answer that well will spend less without slowing down. The ones that don't will keep discovering their AI budgets the same way Uber did, four months in and already gone.

**Also read:** [Scale AI names Francis deSouza as CEO to drive its enterprise and government AI push](https://startupfortune.com/scale-ai-names-francis-desouza-as-ceo-to-drive-its-enterprise-and-government-ai-push/) • [MiniMax and ByteDance ship rival AI video models as China closes the gap on Sora and Veo](https://startupfortune.com/minimax-and-bytedance-ship-rival-ai-video-models-as-china-closes-the-gap-on-sora-and-veo/) • [South Korea's Kospi posts its biggest single-day gain ever after Microsoft earnings calm AI-bubble panic](https://startupfortune.com/south-koreas-kospi-posts-its-biggest-single-day-gain-ever-after-microsoft-earnings-calm-ai-bubble-panic/)
