# Getting a grip on shadow tokens and AI blowouts

> Source: <https://www.cio.com/article/4199599/getting-a-grip-on-shadow-tokens-and-ai-blowouts.html>
> Published: 2026-07-24 12:00:00+00:00

Four months of Claude Code — that’s all it took for Uber to burn through its entire annual budget for AI. Token after token, engineers embraced the platform with few control mechanisms tying costs to outcomes. The result was a budget runaway and [a clear case study](https://www.forbes.com/sites/janakirammsv/2026/05/17/uber-burns-its-2026-ai-budget-in-four-months-on-claude-code/) in how limited oversight snowballs into an AI blowout.

This is a phenomenon I like to call “shadow tokens” — AI credits paid for by the company but largely invisible to decision-makers. Too many engineers have the final say over how much they consume and, therefore, what it costs. This all-you-can-eat attitude is part of the reason why [Microsoft is reportedly](https://www.theverge.com/tech/930447/microsoft-claude-code-discontinued-notepad) winding down many internal licenses across key engineering teams and why [one in five organizations](https://www.thestreet.com/investing/the-next-phase-of-ai-spending-is-already-underway) is missing its AI spend forecast by more than 50%.

And the trend is only accelerating. By 2028, [Gartner predicts](https://www.cio.com/article/4189149/ai-coding-token-costs-are-on-track-to-rival-human-payroll.html) that AI coding costs (driven by this kind of ungoverned consumption) will be as much per developer as the salary companies pay that person.

LLMs and agents introduce a new class of variable cost that scales with behavior rather than headcount, putting enterprises on the hook for tools that balloon with workload. I don’t see this as enterprises overspending because they’re reckless — it’s down to a lack of managerial oversight, budget alignment that demands a proven return on investment, and engineer education on how much is too much.

Going forward, CIOs need to thread the AI needle between governance that encourages transparency and reasonable spend without stifling innovation.

The issue is that AI isn’t a traditional line item. Previously, enterprise leaders onboarded software-as-a-service (SaaS) with a good idea of the total cost. An allocated software seat or annual contract was a known quantity. The cloud added some variation (with fluctuations depending on hosting size), but instances were still modelable. AI flips this status quo on its head — the unit of consumption is behavior and the cost is exponential.

And these specifics aren’t immediately apparent at pilot. Tools can appear inexpensive in controlled experiments yet unpredictably scale depending on session length, context window size, model selection and whether agents run in parallel. This is the fallacy of the $20-per-seat enterprise plan — tokens are charged separately at API rates with no ceiling. The final dollar value of any session is set by factors that finance can’t always model in advance, particularly when these decisions usually rest with the engineers themselves.

According to [Deloitte](https://www.deloitte.com/cz-sk/en/services/consulting/research/the-state-of-ai-in-the-enterprise.html), only 21% of organizations deploying agents have a mature governance model, a real concern because they’re token-eating machines. This is what was happening at Uber — Claude Code in agentic mode was autonomously reading codebases, planning changes across dozens of files and opening pull requests. Each step quickly adds up, with Anthropic’s own documentation noting that agents consume approximately seven times as many tokens as standard sessions.

This is shadow IT and shadow AI, evolved. This time, however, many leaders approved the tool in question without guardrails governing consumption. AI hype adds fuel to the fire and normalizes long sessions. Uber’s CTO, for example, [described](https://x.com/praveenTweets/status/2033627282418655711) a company-wide shift toward “agentic software engineering” with employees “who are quietly experimenting, quietly shipping and quietly pushing things forward”. This is an exciting way to test the limits of what’s possible, certainly, but it’s also a position that goes a long way to explaining how the company spent its annual AI budget by April.

Engineers haven’t done anything wrong here. In fact, they’re adopting and experimenting as instructed, with Uber creating leaderboards and ranking users by token consumption. More use led to a better ranking, reflecting a culture that lauds new ways of doing things. This behavior is known as “[tokenmaxxing](https://www.cio.com/article/4178320/tokenmaxxing-when-ai-adoption-metrics-go-bad.html),” and its principal knock-on effect is shadow tokens — quantity-over-quality processes that leaders struggle to control until they’re fully realized in the budget. Of course, if management treats adoption metrics as performance metrics, then engineers can’t be blamed for using more tokens. The tension is that the teams driving adoption aren’t the ones managing spend.

None of this is meant to dismiss AI’s productivity possibilities and potential return on investment. Developers save [3.6 hours](https://getdx.com/blog/ai-assisted-engineering-q4-impact-report-2025/) per week, achieve 60% higher pull request throughput and cut onboarding time in half with automation. Meanwhile, Uber shared that roughly 11% of live backend updates were written by agents with no human in the loop. However, these wins aren’t the problem — it’s that too many teams aren’t connecting input to output. I’ve spoken to admins who discovered their token spend had tripled in a single quarter after using heavier models or accidentally doubling up on agentic applications. Nobody knew until the financial damage was done.

Automation needs to happen sustainably with an eye on the bottom line. In my view, a much better metric for achieving this is AI yield — the measurable business or engineering output generated per dollar spent on tokens. Otherwise, without a feedback loop, even genuinely productive teams are flying blind.

Creating that throughline between AI investment and token consumption starts with established financial metrics. This is possible via maximum spend limits (dictated by spend tagging, workload tiering and cost-per-output benchmarks) per team or project. Then, any additional allocation requires approval, closing the loop between the engineers spending the tokens and the leaders paying for them. AI isn’t cheap and teams should demonstrate a bang for their buck.

This is something we do with our engineering team at Hexnode. Resource allocation for Claude Code and Cursor is tied directly to ROI rather than letting consumption run open-ended. Given the pay-as-you-go nature of these tools, a firm usage limit per team offers simple but essential control.

Similarly, there’s room to apply some of the governance principles IT uses for device management. Things like policy enforcement, role-based access, real-time monitoring and automated alerts can flag usage behavior in advance. Uncovering such insights at the token layer works to identify power users and prevent excessive spending.

We also need to encourage cultures that praise outputs that actually achieve efficiency. AI applications that result in shipping faster, reducing rework and cutting review cycles are gains that should be celebrated. If your company hosts leaderboards, frame unnecessary token burn as wasteful rather than valuable. The organizations creating healthier consumption habits work with their engineers to understand not just how to use AI, but what responsible use looks like and what it costs.

This is a conversation teams need to have now. Anthropic [just ended flat-rate pricing](https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan) for programmatic workloads from June 15. Now, agents, continuous integration pipelines and automated workflows draw from a dedicated monthly credit pool billed separately from the subscription. Once that pool is exhausted, agent tasks either stop entirely or overflow to extra billing. Work can either get very expensive or grind to a halt for teams that aren’t prepared.

Getting a grip on shadow tokens means better rules and tools connecting spend to outcomes. Only by building the financial and cultural infrastructure that encourages sustainable adoption can leaders see what they’re spending, connect it to what they’re getting and course-correct before the costs become a crisis. Ultimately, shadow tokens are only invisible if we choose not to look.

**This article is published as part of the Foundry Expert Contributor Network.****Want to join?**
