TL; DR: Tokenmaxxing Or Reinventing the Wheel #
Your organization counts AI tokens, seats, and pilots, but can anyone name a single decision those numbers actually changed? Tokenmaxxing is only the symptom; five old Agile Laws explain the cause, and each one comes with a test you can run this week. There is no need to reinvent the wheel with AI transformations and learn the hard way what the veterans of other transformations already figured out.
Thesis: Tokenmaxxing is the vanity metric of pushing low-value work through an AI tool solely to inflate usage metrics. Tokenmaxxing emerged in 2026, when large technology companies began ranking employees by token consumption on internal leaderboards. The behavior is rational for the individual but useless for the organization because tokens measure input rather than outcomes. The five Agile Laws in this article explain why organizations keep making this mistake and what to measure instead.
Disclaimer: I belong to those who read Charniak/McDermott’s book on “Artificial Intelligence” decades ago. Of course, I make use of AI, for example, for research, proofreading, challenging story arcs and article structures, or summarization. It is a tool, and a powerful one… 🤷♂️
🎓 🇬🇧 The A3 Delegation System Founding Workshop — September 28-29, 2026
Your team already delegates work to AI: reports, research, customer feedback analysis, stakeholder communication, or parts of operational workflows.
But can you answer these questions without improvising?
- What may AI decide, and what must remain a human decision?
- What does “good enough” mean for this particular work?
- Who verifies the result before somebody acts on it?
- Who checks whether the delegation still works after the model or workflow changes?
If those answers live in one person’s head, or nowhere, your problem is no longer prompting. You have a delegation problem. The A3 Delegation System gives you a practical way to decide what AI may do, hand over the work clearly, define acceptable results, and inspect the delegation over time.
During two hands-on sessions, you will apply the system to a workflow. You will leave with a clear understanding of how to apply the A3 Delegation System to your workflows so that team members or stakeholders can understand, challenge, and continue your AI delegation work. Everything you learn is directly applicable to your situation the next day. The class is in English.
👉 Join the Workshop Now — $199: The A3 Delegation System Founding Workshop — September 28-29, 2026
🗞 Shall I notify you about articles like this one? Awesome! You can sign up here for the ‘Food for Agile Thought’ newsletter and join 35,000-plus subscribers.
🎓 Join Stefan in one of his upcoming training classes!
The Folly of Tokenmaxxing #
In April 2026, Fortune reported, citing The Information, that an employee at Meta had built an internal leaderboard ranking colleagues by how many AI tokens they consumed, drawing on usage from more than 85,000 employees and displaying the top 250. What we now know as "Tokenmaxxing" was born. The highest-ranked user, in Fortune's wording, "averaged 281 billion tokens" across the 30-day window. The leaderboard handed out titles: "Token Legend" and "Cache Wizard." Neither Mark Zuckerberg nor CTO Andrew Bosworth made the top 250.
At Amazon, the same failure mode surfaced as employees reportedly pushing low-value work through the company's agentic tool to inflate their usage. As one employee put it: "Some people are just using MeshClaw to maximise their token usage."
You do not need an AI expert to diagnose that "Tokenmaxxing" is a folly. An economist from 1975 will do the trick. Enter Mr. Goodhart.
My point is that every new, possibly paradigm-shifting technology or framework arrives with acolytes who relearn the hard way what previous generations already figured out. It is correct that AI lowers the marginal cost of producing work. However, it does not lower the cost of choosing the right work, integrating it, or being accountable for the result. Many of the enterprise AI failure modes currently filling your LinkedIn feed follow from that gap, and five old laws describe them well enough that we should stop acting surprised; every veteran of "Agile" can tell you instantly that tokenmaxxing is a folly.
Cannot see the form? Please click here.
Goodhart Predicted the Tokenmaxxing Leaderboard in 1975 #
Charles Goodhart, then an adviser to the Bank of England, wrote that "any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes" in a 1975 conference paper on UK monetary management. He offered it as an aside.
A correction while I am here. The sentence everyone quotes, "When a measure becomes a target, it ceases to be a good measure," is not Goodhart's. Marilyn Strathern, a Cambridge social anthropologist, wrote it in 1997 on page 308 of a paper about audit culture in British universities, offering it as a paraphrase of Goodhart and citing Keith Hoskin. I have repeated the misattribution myself, including in my earlier article on agile laws. My apologies and a hurrah to the capabilities of research agents!
What AI changes: The measure is now generated automatically, continuously, and per person, at a granularity no manager could have collected in 1975. Gergely Orosz reported in April 2026, citing engineers at Salesforce, that internal tools displayed minimum spending targets of $100 for Claude Code and $70 for Cursor, with a Mac widget refreshing every 15 minutes. He also reported an internal Microsoft token leaderboard running since January, where one engineer told him: "I am conscious of not wanting to be seen as 'uses too little AI,' and I'm not ashamed to say I need to do tokenmaxxing to do this."
What AI does not change: People optimize for what is counted. Jake Paul, a product and innovation analyst at Kyle and Co, described the mechanism to SHRM in one line: "The path from informal leaderboard to team OKR to formal competency on a performance review is short." Logan Wolfe, a partner at Kyndryl, told CIO what that produces: "When token usage becomes the KPI, you incentivize output volume over outcomes like efficiency, quality, and risk reduction."
Satya Nadella got the relationship right in Microsoft's FY26 Q4 earnings release on July 29, 2026, when he said the company is "ensuring every customer can turn tokens into business results." Tokens are the input. Business results are the outcome. Goodhart enters at the point where an organization measures the first and quietly assumes the second. As we know, that assumption has failed in every "Agile transformation" before; sending as many people as possible to Scrum Master training classes does not automatically improve an organization's bottom line.
Your test: For every AI usage number your organization reports upward, ask the person who reports it which decision the number informs. If nobody can name one, the number is AI adoption theater.
The Payoff Arrives After the Redesign #
Robert Solow named the pattern in 1987, reviewing a book for the New York Times: "You can see the computer age everywhere but in the productivity statistics."
Erik Brynjolfsson, Daniel Rock, and Chad Syverson explain the mechanism. General purpose technologies "enable and require significant complementary investments, including co-invention of new processes, products, business models and human capital." The organization has to invent the new work before the new tool pays for it, and during the invention phase the numbers look worse rather than better. (That is the famous "J curve" during adoption phases.)
Microsoft's own randomized trial of more than 6,000 workers across 56 firms shows the tool's early task-level effects, before any documented system-level redesign. Regular Copilot users reduced their weekly time reading email by half an hour, an 18% cut, replied nearly 10% faster (46 minutes), and finished documents nearly a full day sooner. They also created about 11% more Word documents and read 14% more of their colleagues' documents. Meeting time did not fall. The authors state plainly that they "do not observe anything about the content produced by these workers," so quality is unmeasured. (All four authors work at Microsoft, which is worth knowing when reading the interpretation.)
Read those results together rather than as a scorecard. They are consistent with a rebound effect: as producing a document got cheaper, the same workers produced more documents while also consuming more of their colleagues' work. The study does not prove that every saved minute created fresh demand. It does show why task-level efficiency cannot be treated as system-level capacity without checking.
Coming back to the example I mentioned above: An Agile transformation was declared successful when enough people had attended a two-day Scrum Master course. Two days to understand systems, self-management, organizational change, and how to decide what is worth building. The Scrum Master training was real; I put a lot of work into turning it into a worthwhile experience for the participants. However, the organizational redesign that would have made it useful never happened, because nobody asked for it, and, most of the time, nobody lobbied for it or put it at the top of the leadership agenda. (Remember "We Tried Baseball and It Didn't Work" on Ron Jeffries' site?) Copilot licenses are the 2026 equivalent of the Scrum Master class purchase orders back in 2017.
Deloitte's EMEA research from October 2025 shows investment running ahead of returns: 85% of surveyed organizations had increased AI investment over the previous 12 months, 91% planned to increase it again, and 6% reported payback in under a year. That does not prove the investments will fail. My own argument above predicts a lag. It shows why adoption statistics cannot stand in for realized value.
Conway Explains Your AI Islands #
Melvin Conway concluded in 1968 that "organizations which design systems (in the broad sense used here) are constrained to produce designs which are copies of the communication structures of these organizations."
What AI changes: The copies now appear in weeks rather than years, because any department can prototype an assistant without asking anyone.
What AI does not change: The communication structure being copied. Each department that runs its own AI initiative produces its own assistant, and none of those assistants talk to each other, because the departments do not talk to each other. The greenfield prototype has no contact with operational reality because the team that built it often has no reporting line to the people who own the process. "AI islands" is Conway's Law with a 2026 vocabulary.
Where practitioners get stuck: The organization responds by establishing a central AI platform team, which can become a bottleneck for every department that wants to ship anything. Conway's Law offers no escape from itself; it only tells you which structure you are about to reproduce.
Your test: Draw your current or planned agent estate as a diagram. If it matches your org chart, your architecture may be accidental, and it is worth asking who chose it.
Larman Explains Why AI Became a Tool Rollout #
Craig Larman's first law states that "[o]rganizations are implicitly optimized to avoid changing the status quo middle- and first-level manager and 'specialist' positions & power structures." The second predicts that any change initiative gets reduced to redefining or over the new terminology to mean basically the same as the status quo.
What AI changes: Nothing about this. Larman's laws are about political economy, and the political economy of a license purchase is identical to that of a training budget.
What AI does not change: Which changes are permitted. AI adoption gets reduced to tool rollout because tool rollout threatens nobody's position on the org chart. "AI transformation" most often comes to mean "we bought Copilot seats." McKinsey's November 2025 survey found 88% of organizations using AI regularly in at least one function, and 39% able to attribute any EBIT impact to it, most of them below 5%. That gap is consistent with the missing co-invention Brynjolfsson describes: widespread tool adoption typically comes without comparable evidence of workflow and operating-model redesign.
The fourth corollary is the one that should make my own profession uneasy: If managers and specialists are still displaced after the change has itself been changed, they become the coaches and trainers for it. When displaced managers and specialists reappear as prompt-engineering trainers, they occupy the same structural position as the certified agile coaches of 2016. Larman's fifth law is the exit: Culture follows structure.
Your test: Name one consequential decision, handoff, or control point that will change within six months, and name who will own the resulting outcome. Write it down with a date. If neither authority nor workflow constraints move, training and licenses are unlikely to change the culture around the work.
Little Explains the AI Pilot Inventory #
John Little's queuing formula says that in a stable system, average flow time equals average work in progress divided by average throughput.
What AI changes: The cost of starting. A prototype that used to require a budget request now requires an afternoon, so far more work enters the system.
What AI does not change: The throughput of everything downstream. One platform team, one security review, one procurement path, one operations group willing to own a new dependency. Release more pilots into a system with fixed exit capacity, and the average lead time rises; it is basic arithmetic.
One caution against reading the failure statistics too fast, though: S&P Global research, reported by CIO Dive in March 2025, found that the average organization scrapped 46% of AI proofs of concept before they reached production. A high kill rate can be a sign of health. Killing something quickly is what good experimentation looks like, and the number tells you nothing on its own.
Your test: Count three things: Active pilots, transitions to production in the past quarter, and median days from start to a scale-or-stop decision. A large, aging inventory with no decisions attached is a common AI adoption failure mode. A fast kill rate, on the other hand, is not.
Brooks Moves, He Does Not Leave #
Fred Brooks wrote in "The Mythical Man-Month" in 1975 that "adding manpower to a late software project makes it later." His mechanism was ramp-up time plus communication overhead, since the number of pairwise links grows faster than the headcount does.
What AI changes: The social half of that cost. Agents remove most of the onboarding delay. They have no careers to build and no territory to defend, so nobody has to be persuaded that an eleventh agent belongs on the project.
What AI does not change: The rest of it, which is the expensive part, opportunity costs considered. For example, context is fast. Producing context that is correct, current, and sufficiently bounded is not; consider, for example, data quality. Parallel agents still require task decomposition, shared architectural constraints, state management, integration, and testing, and their outputs still embody architectural choices and assumptions. They do not argue with each other unless specifically instructed, which is another can of worms. Their silence sounds like a saving until two of them proceed from incompatible assumptions and nobody notices until integration. Code generation scales faster than coherent integration.
Every consequential output needs a credible verification path. Tests, rules, and automated comparison cover a good deal of it. For ambiguous or high-impact work, that path still ends in qualified human judgment, and buying more compute does not expand that capacity. How far parallel agents scale before integration cost dominates is unsettled, and I have seen no credible data either way.
So Brooks is not repealed, but the unit of analysis moves: Adding agents helps when work can be decomposed cleanly and verification costs less than code generation. Where those conditions fail, the result is a larger queue of unverified work.
Your test: Measure your verification capacity, in hours per week and in automated checks you actually trust, alongside your generation capacity in tokens. Then check which of the two you have been buying.
The Tokenmaxxing Folly Conclusion: Productivity and Transformation Are Different Purchases #
If your organization switched off every AI tool tomorrow, what would change? Most people answer with tasks: writing would take longer, slides would take longer. That is a legitimate answer, as faster writing is a real gain, and the Microsoft trial measured it. However, if that is the whole answer, you bought productivity tooling. Take the gain, but do not call it an "AI transformation." Transformation starts when workflows, decision rights, service levels, or operating economics change in a way you can measure, and skipping that step does not exempt you from any of the five laws above. All you did was add a new, possibly paradigm-shifting technology to a legacy organization.
So avoid that failure by deciding how processes and the organization need to change to maximize the new technology's potential. And that is where the second purchase starts.
Building that judgment is what I designed the AI4Agile Online Course around: applying product judgment to AI adoption rather than treating AI as a subject that stands apart from how value gets built.
Key Questions This Article on the Folly of Tokenmaxxing Answers #
What Is Tokenmaxxing?
Tokenmaxxing is the practice of running low-value work through an AI tool to inflate a usage metric. It surfaced in 2026, when companies including Meta, Microsoft, and Salesforce began tracking employee token consumption on internal leaderboards. Amazon employees reported pushing junk tasks through the company's agentic tool solely to raise their ranking, a behavior they themselves dubbed.
Why Is Token Usage a Bad Measure of AI Adoption?
Token usage counts input, not outcome. Goodhart's Law predicts that any statistical regularity collapses once someone applies pressure to it for control purposes, and an tokenmaxxing internal leaderboard turns consumption into a target employees can game. The useful test is whether anyone can name a decision the number informs. If nobody can, the metric exists to demonstrate progress.
Does Brooks's Law Still Apply to AI Agents?
Partly. Agents remove the social onboarding cost Brooks described, because they have no careers to build and no territory to defend. The remaining costs are untouched: task decomposition, shared architectural constraints, integration, and verification. Adding agents helps only when the work is cleanly decomposed and verification costs less than generation. Otherwise the result is a larger queue of unverified work.
Why Do So Many Enterprise AI Pilots Never Reach Production?
Little's Law explains it: AI cuts the cost of starting a pilot to an afternoon, while downstream throughput stays fixed at one platform team, one security review, and one procurement path. More work enters than exits, so average lead time rises. S&P Global research found the average organization scrapped 46% of AI proofs of concept before production.
What Is the Difference Between AI Productivity Gains and AI Transformation?
Productivity means tasks take less time. Transformation means that workflows, decision rights, service levels, or operating economics change measurably. Microsoft's own randomized trial found that regular Copilot users cut half an hour a week from their email reading time. At the same time, meeting time did not fall, which is a productivity gain within an unchanged workflow.
Folly of Tokenmaxxing — Related Articles #
[Agile Laws: From Conway to Goodhart to Parkinson to Occam’s Razor](https://age-of-product.com/agile-laws-distributed-teams/).
[Product Washing: The Pitfalls of a Superficial Product Operating Model Transformation](https://age-of-product.com/product-washing/).
[Why Leaders Believe the Product Operating Model Will Succeed Where Agile Initiatives Failed](https://age-of-product.com/product-operating-model-agile-failure/).
👆 [Stefan Wolpers: The Scrum Anti-Patterns Guide](https://geni.us/wtS2aa) (Amazon advertisement.)
📅 Training Classes, Workshops, and Events #
Learn more about the Folly of Tokenmaxxing with our AI and Scrum training classes, workshops, and events. You can secure your seat directly by following the corresponding link in the table below:
| Date | Class and Language | City | Price |
|---|---|---|---|
| 🖥 💯 🇬🇧 August 27 to September 17, 2026 | GUARANTEED: | ||
| Live Virtual Cohort |
€499 incl. 19% VAT (If applicable.) |
| 🖥 💯 🇬🇧 Sep 28-29, 2026 | GUARANTEED:
| Live Virtual Class | $199 incl. 19% VAT (If applicable.) | | 🖥 🇩🇪 Sep 30 to Oct 1, 2026 | |
Live Virtual ClassAI4Agile BootCamp #9 (English; Live Virtual Cohort)Live Virtual CohortSee all upcoming classes here.
You can book your seat for the training directly by following the corresponding links to the ticket shop. If your organization's procurement process requires a different purchasing approach, please contact Berlin Product People GmbH directly.
✋ Do Not Miss Out and Learn More about the Folly of Tokenmaxxing — Join the 20,000-plus Strong ‘Hands-on Agile’ Slack Community #
I invite you to join the “Hands-on Agile” Slack Community and enjoy the benefits of a fast-growing, vibrant community of agile practitioners from around the world.
If you would like to join all you have to do now is provide your credentials via this Google form, and I will sign you up. By the way, it’s free.