Everyone knows AI costs are rising. What fewer expected is that efficiency gains themselves may be driving the bill higher.
As improved efficiency lets research labs develop and deploy more powerful, more expensive models, and as users find increasingly sophisticated applications for them, like agentic workflows, token consumption keeps climbing, according to research firm Gartner. That combination is driving up overall inference costs, so much so that Gartner predicts AI inference costs per agentic workflow will increase more than fivefold through 2028.
This highlights an inference paradox, the subject of the report: even as unit economics improve, overall AI costs continue to rise without a clear or predictable path to matching value. The trap is that, while AI agents consume many more tokens, companies are encouraged to invest in these advanced assistants rather than AI chatbots because these agents are more likely to deliver returns and make step changes in the organizations.
Credits: Gartner
Even though those returns aren't predictable or guaranteed at the moment, and need to be higher than ever to justify the cost, the promise is enough to keep enterprises investing in AI solutions. While it would be most beneficial for companies to adopt a take-it-slow approach, outside pressures are unlikely to allow that, according to Scott Bickley, Advisory Fellow at Info-Tech Research Group.
"The current environment has created a top-down fervor, in fact a mandate, for virtually all enterprises to aggressively adopt AI en masse," Bickley told The Deep View. "This blind foray into the AI abyss often lacks the in-depth understanding of the total cost of ownership, can ignore the culture of technology adoption within a given enterprise, and makes it difficult for one to advocate for anything but an 'innovator/early adopter' position."
This demand is so insatiable that it has become the primary focus and biggest revenue generator for many leading AI labs, including OpenAI and Anthropic. As a result, these labs are also scrambling to find enough compute to meet demand. For instance, on Monday, news broke that Nvidia would provide up to $105 billion in financing for OpenAI's data center in Ohio. Nvidia highlighted in its blog post that AI factories are the "defining infrastructure" of the AI era, that "compute is revenue," and, as a result, it sees its responsibility to help secure those resources.
Our Deeper View #
Typically, in any field, increases in innovation are regarded as positive. That's actually the beauty of technology: further developments lead to discoveries that would never have been possible without the previous advancements, creating a virtuous cycle. Yet with AI, it's a bit of a different story. Since AI was widely adopted, compute constraints have emerged, and they have only been exacerbated by growth in research on the development side. Compute and inference costs haven't kept up because both are based on finite, real resources. So, in a rare case for innovation, it may actually be wiser and more beneficial to , an idea supported by some of the biggest companies and experts in AI.