August 10, 2026, (Inside AI) — The cost for enterprises to run AI models has plummeted to its lowest level this year, with average inference prices hovering between $1.16 and $1.18 per million tokens from August 6 to 8. That marks a steep decline from $2.04 on May 31 and $1.45 in late July, according to new data from investment bank Jefferies.
The figures, sourced from US research firm Silicon Data, track pricing across business API providers and open-weight inference platforms used by software developers. The drop signals a fundamental shift in the economics of enterprise AI, driven by an intensifying global price war and the rapid adoption of low-cost Chinese open-source tools, including those from DeepSeek.
Jefferies analysts, led by Thomas Chong, attributed the trend to an “increasing emphasis on cost efficiencies” across both the US and Chinese tech ecosystems. The convergence of competitive pricing and open-source innovation is reshaping how businesses access and deploy large language models.
Price War Intensifies as Open-Source Models Disrupt #
The cost collapse reflects a broader industry realignment. Proprietary API providers like OpenAI and Anthropic have been forced to cut prices aggressively in response to open-weight alternatives that offer comparable performance at a fraction of the cost. DeepSeek’s models, for instance, have gained traction by delivering state-of-the-art reasoning capabilities without the licensing fees of closed systems.
Silicon Data’s index reveals that open-source inference platforms now account for a growing share of enterprise usage, particularly in price-sensitive markets. The availability of models like Llama 3 and Qwen 2 on cloud infrastructure has enabled developers to bypass traditional API gateways, further commoditizing inference.
“The price decline coincided with an increasing emphasis on cost efficiencies across both the US and Chinese tech ecosystems,” the Jefferies analysts noted. This dual pressure from supply-side competition and demand-side optimization suggests that prices may fall further before stabilizing.
Chinese Open-Source Push Redefines Global Pricing #
Chinese firms are leading the charge in affordable computing, challenging the dominance of US-based providers. DeepSeek’s recent releases have demonstrated that cutting-edge AI need not come with a premium price tag, forcing incumbents to rethink their monetization strategies. The trend mirrors the commoditization seen in cloud computing, where scale and efficiency eventually trumped brand loyalty.
Industry observers note that the price war benefits enterprises in the short term but raises questions about the sustainability of AI business models. If inference becomes a low-margin utility, the burden of differentiation will shift to application-layer services and custom fine-tuning. For now, however, the savings are real and measurable.
Silicon Data’s methodology captures a representative sample of API and open-weight pricing, weighted by usage patterns observed across thousands of enterprise deployments. The index excludes one-time promotional discounts, focusing on sustained list prices that reflect true market conditions.
The decline also reflects technical improvements in model efficiency. Advances in quantization, distillation, and hardware-aware optimization have reduced the computational cost per query, enabling providers to maintain margins even as sticker prices drop. This virtuous cycle of cheaper inference and broader adoption is accelerating the integration of AI into everyday business workflows.
Yet the race to the bottom carries risks. Smaller AI startups without deep pockets may struggle to compete, leading to consolidation around a handful of well-capitalized players. Regulatory scrutiny could also intensify if predatory pricing practices emerge, though no such allegations have been made public.
For enterprises, the message is clear: the cost of intelligence is falling faster than expected, and the window to lock in long-term contracts at favorable rates may be narrowing. As the market matures, the focus will shift from raw model access to the quality of integrated solutions and data governance frameworks.