cd /news/artificial-intelligence/us-labs-cut-ai-inference-costs-nearl… · home topics artificial-intelligence article
[ARTICLE · art-98976] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

US labs cut AI inference costs nearly 25% amid price war

Average prices for AI inference from leading US labs dropped nearly 25% between mid-July and mid-August, according to Silicon Data analysis reported by the Financial Times. OpenAI cut prices on GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% on July 30, while Anthropic launched Claude Opus 5 at half the price of its predecessor, Fable 5, amid pressure from Chinese providers DeepSeek and Moonshot. GPT-4-class performance, which cost over $20 per million tokens in late 2022, now costs less than $1 per million tokens by mid-2026, a decline of more than 95% in roughly three and a half years.

read2 min views2 publishedAug 16, 2026
US labs cut AI inference costs nearly 25% amid price war
Image: Cryptobriefing (auto-discovered)

Photo: BalticServers.com / Wikimedia Commons / CC BY-SA 3.0 (https://creativecommons.org/licenses/by-sa/3.0) OpenAI and Anthropic slash prices on mid-tier models as Chinese competitors force a reckoning on AI profit margins

The cost of thinking, at least the artificial kind, just got a lot cheaper. Average prices for AI inference from leading US labs dropped nearly 25% between mid-July and mid-August, according to Silicon Data analysis reported by the Financial Times.

Who cut what, and by how much #

OpenAI moved first on July 30, taking a machete to pricing on two of its three GPT-5.6 model tiers. GPT-5.6 Luna, the company’s mid-range offering, saw an 80% price reduction. GPT-5.6 Terra got a 20% haircut. The flagship GPT-5.6 Sol held steady, though OpenAI sweetened the deal by accelerating its performance options at the same price point.

Anthropic followed a similar playbook, launching Claude Opus 5 at half the price of its predecessor, Fable 5.

The DeepSeek effect #

The pressure is coming primarily from Chinese providers, with DeepSeek and Moonshot leading the charge. These firms offer models that either match or closely approach US performance benchmarks, but at dramatically lower price points.

To put the broader trajectory in perspective: GPT-4-class performance cost over $20 per million tokens in late 2022. By mid-2026, that same tier of capability costs less than $1 per million tokens across various providers. That’s a decline of more than 95% in roughly three and a half years.

Good for users, complicated for investors #

For enterprises that have been cautiously experimenting with AI, cheaper inference is unambiguously good news. Lower costs remove one of the biggest barriers to broader adoption, letting companies run more experiments, deploy more agents, and integrate AI into workflows that previously didn’t pencil out at higher price points. US AI labs have attracted enormous valuations built partly on the expectation that inference would remain a high-margin business. When your revenue-per-query drops by 25% in a month, the financial models that justified those valuations need revision. OpenAI’s decision to hold GPT-5.6 Sol’s price while cutting Luna and Terra suggests it’s betting on a tiered strategy, where the best model keeps its margins while cheaper tiers fight for volume.

US labs have spent billions building out inference infrastructure. Those investments were justified by revenue projections that assumed certain price levels. When prices fall faster than usage grows, the return on that infrastructure spending gets stretched, potentially reshaping how investors evaluate the entire AI supply chain from chips to cloud providers to the labs themselves.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/us-labs-cut-ai-infer…] indexed:0 read:2min 2026-08-16 ·