# Button-mashing never worked in StarCraft. Tokenmaxxing won’t work in AI

> Source: <https://fortune.com/2026/07/29/button-mashing-gaming-tokenmaxxing-nnamdi-iregbulem-lightspeed-ai/>
> Published: 2026-07-29 12:30:00+00:00

In the race to implement and deploy AI in our organizations, we risk making the same mistakes so many gamers do in their battles for the top spot. I say that with some authority: I’ve always been a gamer. Real-time strategy games like *Age of Empires*, *StarCraft*, and *Command & Conquer*; online battle arenas like *DotA* and *League of Legends*. I had a competitive streak, too, pouring countless hours into mastering their intricacies and sharpening my strategies. After all those late nights, the analogies between gaming and the emerging world of AI are hard to ignore.

In years past, a common focal point for gamers trying to elevate their skills was actions per minute or APM: the pace of actions executed by the player. This was often seen as a proxy for player skill, as more advanced players could better handle the cognitive rush of commanding dozens of units across multiple fronts, presumably leading to more effective play, eventually overwhelming the opponent.

This metric became so popular that it became a target in and of itself. Novice players would look to the recorded APMs of expert-level players and set that as a goal to be achieved. And that’s exactly when it lost its value, a perfect example of [Goodhart’s Law](https://urldefense.proofpoint.com/v2/url?u=https-3A__www.splunk.com_en-5Fus_blog_learn_goodharts-2Dlaw.html&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=cpmHBx4N3t79vyalERF0qdOUEXoKY30yf69ACSdDUXg&e=): when a measure becomes a target, it ceases to be a good measure. Merely pumping up one’s APM by frantically clicking around the game did nothing to help you win, but could certainly generate the illusion you were improving.

Studies corroborate this. In 2022, researchers published a [study](https://urldefense.proofpoint.com/v2/url?u=https-3A__journals.plos.org_plosone_article-3Fid-3D10.1371_journal.pone.0265526&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=i5aGhYAP1htqwEzz6YPVSUBNkY81N-cdsbAsDYdaX-c&e=) examining the behaviors of *StarCraft* players of varying skill levels completing tasks of varying difficulty in the game. While the sample was admittedly small, they found no statistically significant difference in APM between expert and novice players. [An earlier 2014 analysis](https://urldefense.proofpoint.com/v2/url?u=https-3A__journals.plos.org_plosone_article-3Fid-3D10.1371_journal.pone.0094215&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=L0kADLWvChFAX3JRsb6hP9wqYHZZssrf_1YN7zy1o40&e=) of *StarCraft* gamers found that although action and reaction times tended to decline after the age of 24, older players more than made up for this via better judgement and more efficient actions.

The futility of APM extends beyond human operators: when training [AlphaStar](https://urldefense.proofpoint.com/v2/url?u=https-3A__deepmind.google_blog_alphastar-2Dmastering-2Dthe-2Dreal-2Dtime-2Dstrategy-2Dgame-2Dstarcraft-2Dii_&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=B5Ch1JxOTARsKtyU5VulwDXt7UuOCr5mxovoNtJ-0Vo&e=), Google DeepMind’s *StarCraft*-playing AI system, the researchers found the AI performed worse when allowed to click more, as it wasted effort and computation on minutiae instead of grand strategy. In its demonstration games against professional *StarCraft* players, AlphaStar averaged substantially lower APM than the humans and yet emerged victorious nonetheless.

While gaming might have settled the argument some time ago, we’re repeating the same mistakes in AI today. The new obsession? [Tokenmaxxing](https://urldefense.proofpoint.com/v2/url?u=https-3A__blog.pragmaticengineer.com_the-2Dpulse-2Dtokenmaxxing-2Das-2Da-2Dweird-2Dnew-2Dtrend_&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=JeK5R6F6MxL-deqcljoHaeEWhT1SYfUqNwDS3Zz5EWI&e=) – the idea that more tokens is necessarily better, so just let it rip. Agents are all the rage these days, and they are extremely token-consumptive. In fact, Anthropic [has found](https://urldefense.proofpoint.com/v2/url?u=https-3A__www.anthropic.com_engineering_multi-2Dagent-2Dresearch-2Dsystem&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=8OjvjCm3u6Zk4ClGq5TMnFAhpohNbDOCwHQqrzS1eBo&e=) that agents typically consume 4X more tokens than chat interactions, while multi-agent systems consume 15x more than standard chat.

The analogy to gaming is uncanny: users and even companies brag about their token usage in nearly identical ways to gamers bragging about their APM. But wasteful tokens are likely wasteful clicks: a vanity metric, at best.

It’s arguably even worse in the AI context, because all those generated outputs need to eventually be verified and validated. In gaming, you eventually win or lose the match, the natural ultimate arbiter of skill. How does one even measure skillful token usage? It’s not so obvious, and organizations everywhere are actively grappling with this emerging conundrum.

The solution? Focus less on actions or tokens and more on observability and verification. APM, tokens, lines of code, hours logged: each tempts us precisely because it is easy to track. But what we actually care about, useful output, sits one step removed, downstream of judgement and gated by verification. To quote a recent [economic theory paper](https://urldefense.proofpoint.com/v2/url?u=https-3A__arxiv.org_pdf_2602.20946&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=AJ33UTrnWxXNHvhztWzsjMG9Pk4Qc-lGtIp0cR8kqng&e=) on the impact of AI from researchers at MIT, Washington University, and UCLA, “Some Simple Economics of AGI”: “the binding constraint on growth is no longer intelligence. It is human verification bandwidth.”

A helpful framework: the [OODA loop](https://urldefense.proofpoint.com/v2/url?u=https-3A__fs.blog_ooda-2Dloop_&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=h_MW4Ma0fWz5WeUB4rAYDiJ2EeMFAdhIxcXjb0XEYqc&e=) was conceived by U.S. Air Force Colonel John Boyd in the 1970s as a way to conceptualize decision-making during combat operations. Observe the environment, orient yourself in that environment by interpreting the incoming information, decide the best course of action based on your best judgement, and only then act. Boyd’s insight was that whoever observes and orients better gets inside their opponent’s decision cycle and dictates the fight. Faster action alone doesn’t cut it.

Again, we find evidence of the importance of observation and orientation from the video game world. That same study that found no relationship between skill and actions per minute among StarCraft players did find a more fruitful distinguishing characteristic: how and where they placed their attention. Expert players were better observers, the gaze of their eyes covered a wider portion of the screen and they moved their eyes more often, in effect collecting more observations than novice players. Google’s AlphaStar managed its attention by “switching context” about 30 times per minute. Google concluded its success was more due to “superior macro and micro-strategic decision-making, rather than superior click-rate [or] faster reaction times.”

In games, the best players, whether humans or AI, first develop the skill of seeing the whole game board. With that contextual information, strong players build capacity ahead of need, else their ability to grow their armies becomes “capped.” Likewise, in AI the binding constraint is your capacity to monitor, review, and validate the stream of agent work. Strong AI operators build robust and scalable verification ahead of generation in order to avoid costly bottlenecks. [A 2025 analysis](https://urldefense.proofpoint.com/v2/url?u=https-3A__arxiv.org_pdf_2507.09089&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=H4ezm6uaSbqtER0qU3bmTuJbSr4Ljh2ut6XYU-dq38A&e=) of the impact of AI on software development found that engineers spent more than double the time reviewing and modifying AI-generated code (9% of coding time) than waiting for AI code generations (4%). In [a similarly-timed talk](https://urldefense.proofpoint.com/v2/url?u=https-3A__www.youtube.com_watch-3Fv-3DLCEmiRjPEtQ&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=nKqNRu4OIsqKWJSXhwuDB5yKjtCJig3aeu29RpzjtDk&e=) the famous AI researcher Andrej Karpathy argued the bottleneck moved from typing code to verifying code, observing “I’m still the bottleneck, right? Even though that 1,000 lines [of code] come out instantly, I have to make sure that this thing is not introducing bugs and that it’s doing the right thing.”

By the way, competitive gamers eventually found a better yardstick, with Blizzard, the company behind *StarCraft*, eventually adopting “[effective APM](https://urldefense.proofpoint.com/v2/url?u=https-3A__starcraft.fandom.com_wiki_Actions-5Fper-5Fminute&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=hGl4NpU8GkO5WQWXXmxCLb3mSglnm2FLd1Su5-Fgpt4&e=)” as a better metric, which filters out repetitive or redundant inputs, isolating genuine strategic execution. AI unfortunately has no equivalent, a measure of effective, verified agentic output as distinct from raw tokens burned.

Gabriel Petersson, formerly a researcher on OpenAI’s Sora video model team, recently [tweeted](https://urldefense.proofpoint.com/v2/url?u=https-3A__x.com_gabriel1_status_2068483683321798724&d=DwMGaQ&c=euGZstcaTDllvimEN8b7jXrwqOf-v5A_CdpgnVfiiMM&r=r_DP33r3A6LLm8PkSzGOjZ1zLi4__q__rViKSI0dghs&m=j2Pq3It4q6zdxh7REGxQ4N86j5r4NWlcmTpuXMubuTWtS9hfkVaLVLv1rW3REXgk&s=jSFcvRLuTULylXEAoQPOx3ivkOlPqN56jxliaZqfw90&e=) that becoming a top player in these games requires more skill than the vast majority of engineering jobs. I tend to agree. But in the same way that lines of code is not a reliable measure of engineering productivity and APM is a noisy proxy for gaming prowess, token use alone does not imply or guarantee useful, productive output. Optimize the countable proxy and you optimize the wrong thing, often at the direct expense of the right one.

*The opinions expressed in Fortune.com commentary pieces are solely the views of their authors and do not necessarily reflect the opinions and beliefs of *Fortune*.*
