Steve Yegge’s recent essay, “The Flat Curve Society”, makes an argument that should change how we think about AI spend. Commodity intelligence, he writes, “will soon stop growing exponentially, or at least, it will appear that way”, and “the curve is flattening for most of us”. Once raw capability stops being the main difference between teams, the more important question becomes how well they use it.
Yegge frames the shift as two culture problems in sequence. First, you teach people to spend tokens. Then “your second culture problem emerges, which is teaching people how NOT to spend tokens.” His own line points toward the next stage:
Token spend only signals literacy on the way up… then it flips, and the thing you need to start measuring is token waste.
I think that is right, but I want to push one step further. Token consumption on its own, whether spent well or wasted, is not very informative. It is like CPU hours, an AWS invoice, or a monthly electricity bill. No one runs a business to minimise infrastructure spend in isolation. They care about what that spend enabled. That gap between what organisations are starting to measure, token consumption, and what they actually care about, delivery, is the subject of this post.
We have been here before, one layer down. Early cloud teams often focused on the inputs: server utilisation, CPU hours, storage volumes, reserved-instance coverage, and monthly infrastructure bills. For a while, those numbers dominated the conversation. Teams could sometimes explain their utilisation rate more precisely than they could explain whether the business was getting good value from the spend.
Over time, the question changed. It became less “how much infrastructure are we using?” and more “what are we delivering with it, and at what cost per unit of delivery?”
That shift helped produce FinOps, which connected cloud spend to business value; DORA metrics, which measured software delivery performance through deployment frequency, lead time, change-failure rate, and time to restore; and platform engineering as a discipline focused on improving the path from code to production.
The common thread is that infrastructure is not measured for its own sake. It is measured in relation to the outcomes it supports.
AI is starting down the same path. Right now, many organisations are still in the early stage: they can quote their token bill, but they cannot reliably connect that bill to delivery.
Tokens are the fuel that agentic systems consume. An organisation running agents at scale might report 100 billion tokens a month, £500k of inference spend, a mix of models, and thousands of agent runs. Those are large numbers, but on their own they do not tell us whether the money was well spent.
The problem is easiest to see through a simple comparison. Imagine two companies with identical token bills. One ships two hundred features that customers use. The other generates millions of lines of code that are reviewed, distrusted, and abandoned. The spend is the same, but the outcomes are completely different. The token number is blind to that difference.
Those figures are illustrative rather than measured. The point is not the exact values. The point is that a cost number that looks the same whether the outcome was useful or not is not yet doing much management work.
Token spend is an input. To become useful, it has to be connected to a result.
The past year of enthusiasm is now meeting its first correction. A recent report on corporate AI spending describes companies that spent much of the last year telling employees to use as much AI as possible: Claude Code, ChatGPT, one tool per task, senior people running several sessions at once and building workflows around them. Now they are finding the bill grew faster than the output. In a survey of 300 executives, 68% said they overspent their AI budget over the past year. Gusto’s CFO reportedly discovered he was one of the company’s heaviest AI spenders, with usage costs well above the value of the work it produced.
The response so far has mostly been a cap. Uber reportedly limited employee AI spend to $1,500 a month; Tesla to $200 a week. Moonbounce, a startup whose founder had been telling staff to “use more AI”, found spending climbing fast, engineers struggling to review the volume of AI-generated code, and more than 5% of usage unrelated to work at all: trip planning, chatting with the model for its own sake. It introduced rules against off-task use and cut its bill by roughly 10%, which the founder put at tens of thousands of dollars.
Some of that is straightforwardly good. Recreational token spend is waste by any definition, and removing it is the easy first pass. But look at the denominator in every one of these moves. Uber’s $1,500, Tesla’s $200, Moonbounce’s 10% cut are all measured in spend. The cap limits the input. It says nothing about what the input bought.
That is the same mistake one layer up. A raw spend cap is the token-era version of telling a cloud team to hold CPU hours under a line, regardless of what those CPU hours served. It controls the bill, but it controls it blind. A cap set below what a genuinely valuable workflow needs will quietly destroy delivery no one measured, which is what shows up at the far end of these stories: companies that pulled advanced tools or plugins found employees’ workflows breaking and people falling back to slower manual work. The cap saved money and cost delivery, and because delivery was never on the ledger, only one side of that trade was visible.
So the pullback is real, and the instinct behind it is correct: not all of this spend is buying value. The tool is just too crude. A cap on raw consumption can only answer “are we spending less?”, when the question worth asking is “are we spending well?”. Those two come apart the moment the cheap workflow and the valuable one carry the same price tag.
The idea I would propose, by analogy with FinOps, is TokenOps: the practice of connecting AI consumption to delivery, so that spend can be attributed to the work it supported.
The core mechanism is attribution. Each work item, whether it is a feature, bug, support case, migration, analysis task, or operational workflow, should accumulate a record of what it consumed and what it produced:
With that record, the question changes. Instead of asking “how many tokens did we use?”, we can ask “what came from the tokens this piece of work consumed?”
That is a more useful question for engineering leaders, product leaders, and finance teams, because it has an outcome on the other side of it.
I want to be clear that TokenOps is a proposal, not an existing standard. The analogy with FinOps is strong, but the tooling to attribute tokens to delivery at this level is still immature. I am describing where I think the practice needs to go, not reporting on a solved problem.
Most organisations today, if they measure AI usage at all, measure token spend. The layers that make the number meaningful are usually missing.
A more useful way to think about it is as a stack, read from the bottom up.
Tokens sit at the bottom of the stack. They are infrastructure. Agent runs sit above them, because agents are the immediate consumers of the tokens. Deliveries sit above agent runs, because delivery is the first unit of work that people outside the engineering organisation usually recognise. Business outcomes sit above delivery, because that is ultimately what the organisation cares about.
Measuring token spend without the layers above it is like reporting a fuel bill without knowing where the vehicles went, what they carried, or whether they arrived.
The work of the next few years is not just improving token reporting. It is building the attribution layers above the token bill.
If this direction is right, the useful metrics will look less like price per token and more like delivery per token. Possible examples include: These are guesses at the shape of the metrics, not a settled list. The useful ratios will be discovered through practice, just as software delivery teams discovered which measures were actually helpful and which ones became dashboard noise.
The direction, though, seems clear. The AI-era equivalent of delivery performance measurement will not be built on consumption alone. It will be built on delivery normalised by consumption.
The strongest teams will not necessarily be the ones spending the fewest tokens or the most tokens. They will be the ones that can show the best relationship between token consumption and useful delivery.
Here is a more speculative idea, and I will label it as such.
Imagine each epic or project as a trace that carries the work’s full history: the people involved, the agents used, the models called, the tokens consumed, the tool calls made, the retries, the cycle time, and the outcome. The trace is the unit you follow, in the same way a distributed trace follows a request across services.
One possible metaphor is a delivery convoy. Each convoy carries fuel, crew, tools, and cargo toward a delivery. The question is not just how much fuel it burned, but what it brought back.
That metaphor is close to the world Yegge is already describing. His Gas Town uses the language of fuel, towns, rigs, and worker agents. Tokens are the fuel. Workspaces are places where work happens. Agents consume resources to move work forward. Seen from the accounting side, a delivery convoy is simply a way of asking what a particular body of work consumed and what it delivered.
Gas Town also shows why this accounting layer matters. It is not a single user making a few model calls. As described, it can throw twenty or thirty Claude Code agents at one problem, with those agents organised into roles and watched by other agents whose own cost scales alongside the workers. At that point, the bill is no longer easy to interpret from a model invoice alone.
Saying “team X spent £40k on Claude” does not tell you whether the money bought useful engineering progress, duplicated effort, supervision overhead, failed attempts, or genuine delivery. The useful unit of accounting has to move closer to the work itself: which epic the agents were working on, how many runs it took, what was accepted, what was discarded, and what outcome the spend supported.
If each piece of work had a trace, teams could ask questions that are difficult today. Why did one epic consume four times the tokens of another and deliver less? Was the problem genuinely harder? Was the task poorly scoped? Did the agent loop unnecessarily? Was the model choice wrong? Was a missing tool forcing the system to compensate with repeated reasoning? Those are debuggable questions. They point toward model selection, workflow design, tool availability, task decomposition, review overhead, or governance.
I do not know that “convoy” is the right abstraction, and I do not know that Gas Town’s vocabulary is the one that will stick. But the orchestration layer that could emit these traces is already being built. What is missing is the ledger that turns those traces into cost per delivery.
Something like it will be needed if organisations want to reason about AI spend at the level of work rather than the level of the invoice.
The stack tells us what to measure. It does not tell us how to operate, and that is the harder half.
FinOps worked because it was not only a dashboard. It became an operating model: a repeating loop of work, owned by a named cross-functional practice, with decisions attached to each stage.
If TokenOps is going to be useful, it needs the same three things: a loop, an owner, and a unit of account. The loop might have three stages:
Attribute: wire every token to a work item, not just to a model, team, or cost centre. Today, much AI billing stops at something like “team X spent £40k on model Y this month.” That is roughly equivalent to a cloud bill with poor tagging. The first job is plumbing: stamp every agent run, tool call, and model request with the work item it served, so spend can roll up to features, bugs, cases, migrations, or other delivery units.
Until this exists, the rest of the loop has weak data.
Evaluate: once spend is attributed, compute delivery per unit of spend and look for where consumption and output come apart. Some of what shows up will be waste: retries, runaway loops, excessive context, agents re-reading the same files, or work that is generated but never accepted. Some of it will be legitimate: hard problems sometimes cost more.
This stage exists to distinguish the two and compare real options. Should the workflow use a cheaper model? A smaller context? A better tool? More human judgement? Should the task be decomposed differently? Should the work be done at all?
Govern: turn the findings into budgets and policies anchored to outcomes rather than raw usage. Instead of saying “the platform team gets £100k of tokens a month”, a delivery-anchored budget might say “this class of epic is normally worth roughly N tokens; beyond that, escalate.” Agents that reliably contribute to delivery get more room. Agents whose delivery per token remains low get redesigned, capped, or retired. This is the difference between the flat caps companies are reaching for now, $1,500 a month or $200 a week, and a budget that knows what the spend is for. A flat cap throttles the valuable workflow and the wasteful one at the same number. A delivery-anchored budget gives the workflow that ships more room and reins in the one that does not.
The unit of account that holds this together is cost per delivery, not cost per token. A token price is a rate. Marginal cost per shipped feature, resolved case, or completed migration is something a product owner and a finance partner can both reason about, because it has an outcome attached.
That change of denominator is most of what separates TokenOps from a token dashboard.
Ownership matters too. The FinOps lesson is that this should be a thin cross-functional practice sitting between engineering, product, and finance. It should own the measurement loop, not centralise every spending decision. The teams doing the work should still own the work. A TokenOps function that becomes an approval queue will be routed around quickly.
I would expect the practice to mature in stages.
This is still a proposed model. The hardest part, per-delivery attribution, depends on tooling that mostly does not exist yet. The ratios will also be noisy before they become trustworthy. But the shape is borrowed deliberately from a discipline that already worked once, one layer down.
The first decade of cloud showed that infrastructure needs an operating model, not just a meter. Agentic AI does not need to relearn that lesson from scratch.
The first decade of cloud taught us that infrastructure metrics are not business metrics. Utilisation, CPU hours, and storage were inputs that teams eventually learned to connect to delivery and outcomes.
The first decade of agentic AI is likely to teach the same lesson, with tokens playing the role that compute used to play.
Yegge is right that tokens are becoming a resource organisations need to understand, and that efficiency will become an important skill. I would add one thing: efficiency only means something when measured against delivery. A team that halves its token spend and also halves its useful output has not necessarily become more efficient. It has just spent less. Cutting the bill by removing recreational use, as Moonbounce did, is a real gain; cutting it by capping the workflows that ship is not, and a spend number on its own cannot tell the two apart.
The important question is not how many tokens an organisation consumed. It is what those tokens helped deliver. The organisations that learn to measure the second thing will have a better operating model than the ones still optimising the first.
Disclaimer: The views and opinions expressed in this article are my own and do not represent those of my employer or any affiliated organizations. The content is based on personal experience and reflection, and should not be taken as professional or academic advice.
Measuring Agentic AI by Delivery, Not Tokens was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.