TL;DR — Key Takeaways
Corporate America is learning to route every AI task to the least expensive model capable of doing it well enough. Frontier labs may remain indispensable while someone else increasingly decides when they get paid.
The AI industry spent the past three years convincing Corporate America that intelligence was priceless. Corporate America has started comparison shopping.
The Wall Street Journal recently put a sharper edge on what has been building across the enterprise AI market. Companies are no longer assuming that every AI task requires the smartest, largest and most expensive model available. They are mixing models, reserving premium frontier systems for difficult reasoning while routing more routine work to smaller, cheaper alternatives—including increasingly capable open-weight models developed in China.
The Journal describes companies “shopping à la carte for their artificial intelligence.” That may sound like a straightforward attempt to control runaway AI bills. It is much more than that. It is the beginning of a fundamental reordering of power across the AI stack.
The emerging discipline has acquired an appropriately financial-sounding name: tokenomics.
Tokenomics asks a brutally practical question: What is the least expensive model capable of completing this particular task well enough?
Not which model is the smartest. Not which company has the most famous CEO, raised the most money or posted the best score on the latest benchmark. Which model can perform this specific piece of work at the required level of accuracy, speed, security and reliability for the lowest acceptable cost?
Once enterprises begin asking that question task by task, no model remains indispensable across the entire workflow.
That is how tokenomics operationalizes what I call the indispensability trap. It is also the central argument of my forthcoming book, The Indispensability Trap.
The trap occurs when a foundational technology becomes indispensable to the economy without its original providers remaining indispensable to every transaction or capturing most of the value created above it. We have seen versions of it with railroads, electricity, telecommunications, fiber networks, cloud infrastructure and other technologies that became essential precisely as they became standardized, interchangeable and subjected to relentless price pressure.
AI models may now be entering the same cycle.
Frontier labs made intelligence valuable enough to use everywhere. The cost of using it everywhere is teaching their customers how to use frontier models less.
Amazon Starts Rationing Claude
Amazon provides an almost perfect example.
According to internal documents reviewed by Business Insider, Amazon has been redesigning Alexa+ to reduce its reliance on Anthropic’s Claude models. The documents reportedly showed AWS costs for Alexa+ on pace to reach approximately $1.7 billion in 2026, nearly three times the previous year and roughly 60% above Amazon’s target cloud cost per monthly active user.
Amazon’s response was not to remove Claude from Alexa+. It was to ration it.
The company reportedly began routing more requests through its own models, increasing the use of cached responses, eliminating redundant inference and using deterministic software for requests that did not require a large language model. Internal plans also called for moving some specialized Alexa functions away from Claude Sonnet and onto Amazon’s own systems. The combined changes were projected to more than quadruple the number of customer transactions each unit of computing capacity could support.
That is tokenomics in production.
Claude may still provide the best answer for Alexa+ when the request requires difficult reasoning, nuanced language or capabilities Amazon’s own models cannot reliably deliver. But Amazon sees no reason to pay Claude prices to turn on a light, retrieve a familiar answer or execute a predictable command.
The delicious part is that Amazon has invested billions of dollars in Anthropic. It has helped finance Claude’s development, made AWS Anthropic’s primary cloud provider and positioned the relationship as one of the most important alliances in AI.
Amazon invested billions to make Anthropic indispensable. Now it is investing more money to make sure Alexa does not have to use Anthropic very often.
Model loyalty ends where token economics begins.
This does not mean the Amazon-Anthropic partnership is unraveling. It means Amazon understands the difference between having strategic access to frontier intelligence and paying frontier prices every time someone asks Alexa about the weather.
Anthropic understands the logic, too. Its own guidance on building effective AI agents recommends routing easy and common questions to smaller, less expensive models while reserving more capable models for hard or unusual requests. That is entirely sensible architecture. It is also the mechanism through which premium models become optional.
Routing may begin within one model family, sending easier work to a Claude Haiku-class model and harder work to Sonnet or Opus. There is no durable reason it must end there. Once an enterprise installs an orchestration layer capable of measuring a request and selecting an appropriate model, that layer can eventually choose among Claude, GPT, Gemini, Grok, DeepSeek, Kimi, Llama and models that have not yet been released.
AWS is already productizing the mechanism. Bedrock Intelligent Prompt Routing predicts which model can deliver the desired response at the lowest cost. AWS says the service can reduce costs by as much as 30% without compromising accuracy. Today, its managed routers select between models within the same family, but the larger architectural direction is unmistakable.
The cloud or orchestration platform—not the model developer—decides when the expensive model deserves to be invoked.
That is the control point.
Agents Make the Economics Harder, Not Easier
Agents intensify the pressure because they transform an AI interaction from a response into a process.
A conventional chatbot might receive a prompt and generate one answer. An agent can plan the work, search for information, choose tools, execute actions, inspect results, repair mistakes and verify the final output. One user request can produce hundreds or thousands of model calls.
McKinsey recently reported that enterprise leaders are already experiencing sticker shock from agentic AI. Ninety-three percent of respondents to one McKinsey survey said they had exceeded their AI budgets, while one-fifth of respondents to its broader 2026 State of AI survey said their organizations had constrained AI use because of operating costs. McKinsey identified expensive models being used for simple tasks as one of the causes.
The economics produce a counterintuitive result. The more useful and autonomous AI becomes, the less rational it is to use the most expensive model for every step.
A frontier model may be worth the premium to formulate a plan, resolve ambiguity or make a consequential decision. It may not be worth that premium to classify a document, look up a record, check a status field or reformat an answer. The agent may need the frontier model’s intelligence, but it does not need that intelligence continuously.
OpenAI’s own enterprise data illustrates the scale problem. The company reported that average reasoning-token consumption per organization increased approximately 320-fold over a 12-month period. That is an extraordinary signal of adoption, but it also guarantees that token cost will become impossible to ignore.
The progression is straightforward. Frontier models make AI useful. Greater use makes token costs visible. Agents multiply those costs. Tokenomics makes differences in price and performance measurable. Routers turn provider selection into a software decision. Smaller, cheaper and open-weight models absorb routine workloads. The orchestration and application layers take control of the workflow and customer relationship.
The frontier model becomes a premium component inside someone else’s platform.
It may be the most important component. That does not mean it processes most of the tokens or captures most of the revenue.
Intelligence Is Becoming a Buyers’ Market
This pressure would exist even if the frontier labs were competing only against one another. Chinese models and open weights make it considerably more powerful.
The 2026 Stanford AI Index found that the performance gap between leading American and Chinese models had effectively closed. As of March, the top U.S. model led the top Chinese model by 2.7%, with the gap remaining in the single digits and the two countries’ models trading places near the top over the preceding year. The leading closed model held a 3.3% advantage over the leading open-weight model.
Benchmarks are imperfect, frequently gamed and rarely capture all the requirements of a production enterprise workload. A 2.7% or 3.3% aggregate difference does not mean every model is interchangeable or equally trustworthy.
It does change the purchasing question.
The relevant question is no longer simply, “Which model is best?” It is, “Is the additional performance worth the additional price for this step?”
A small capability advantage can support premium pricing for the hardest reasoning problems. It is far more difficult to use that advantage to justify premium pricing for billions of routine tokens.
The price trend is already clear. Stanford’s 2025 AI Index found that the cost of obtaining GPT-3.5-level performance dropped from approximately $20 per million tokens in November 2022 to seven cents by October 2024. Yesterday’s scarce and expensive intelligence rapidly becomes today’s inexpensive baseline.
That is a punishing treadmill for the frontier labs. They must invest billions of dollars in training, infrastructure, researchers, energy and distribution to remain at the leading edge. Each breakthrough provides a temporary performance advantage. Then competitors reproduce, optimize, distill or approximate that capability at a fraction of the cost.
The frontier keeps moving, but yesterday’s frontier keeps falling into the commodity layer.
This is where the geopolitical argument becomes uncomfortable.
Washington and the American labs have legitimate concerns about Chinese models. Enterprises must account for data security, privacy, intellectual property exposure, censorship, sovereignty and strategic dependence. A company should not route sensitive work to an untrusted model merely because it is cheap.
But these warnings are becoming louder at precisely the moment inexpensive Chinese models are applying severe pricing pressure to enormously valuable American companies. National security and incumbent economic protection are beginning to occupy the same sentence.
That does not make the security concerns false. It does mean they will increasingly have to withstand economic scrutiny. When the price difference is 10x, 20x or 50x, enterprises will not simply ignore it. They will segment workloads, establish policies and route sensitive tasks differently from routine ones.
Security will become one more input to the router.
Whoever Owns the Router Owns the Decision
The model companies’ larger strategic problem is that value is already accumulating above them.
Menlo Ventures estimated that enterprises spent $37 billion on generative AI in 2025, with $19 billion—more than half—going to the application layer. That is where companies such as Cursor and other AI applications own the interface, the workflow, the proprietary context and the customer relationship.
The model makes the experience possible, but the application decides which model receives the work.
G2’s decision this year to establish AI Gateways as a distinct software category provides another important signal. These gateways sit between enterprise applications and model providers, centralizing multi-model routing, failover, semantic caching, token accounting, API-key management, governance and cost control.
Once a gateway sits between the application and the model, switching suppliers can become a configuration decision rather than an application rewrite.
This is the same architectural abstraction that transformed other parts of the technology stack. Standard interfaces make adoption easier, which expands the market. They also make the underlying supplier easier to replace.
The frontier labs therefore face an awkward tradeoff. They need models to become easy to integrate everywhere, but every layer that simplifies integration can also weaken their control over the customer. They need enterprises to consume enormous numbers of tokens, but enormous consumption gives enterprises an overwhelming incentive to optimize token cost. They need agents to make AI essential to business operations, but agents generate so many calls that customers cannot afford to send every one to the most capable model.
Adoption creates the incentive for substitution.
None of this predicts the imminent collapse of OpenAI, Anthropic or the other frontier labs. Frontier capability still matters. Trust matters. Safety, reliability, distribution and proprietary performance matter. There will remain tasks for which using anything less than the best available model is a false economy.
The argument is subtler and more consequential.
Technological indispensability does not guarantee pricing power, transaction ownership or control of the customer.
OpenAI and Anthropic may remain indispensable to the AI economy while becoming optional for most of the tokens the economy consumes. They can build the models that define what is possible while applications, gateways and orchestration platforms decide how often those models are used.
They may provide the most important intelligence in a workflow without capturing most of its economic value.
That is the indispensability trap.
The AI labs wanted tokens to become a new currency of business. They may now be discovering what happens to everyone who sells an increasingly standardized resource priced by the unit.
The frontier labs made intelligence abundant. Tokenomics is how their customers turn that abundance into leverage.
The indispensability trap does not require OpenAI or Anthropic to lose. It only requires them to remain indispensable while someone else decides when they get paid.