cd /news/ai-agents/the-three-waves-of-ai-consumption · home topics ai-agents article
[ARTICLE · art-126380] src=tomtunguz.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Three Waves of AI Consumption

Agent token consumption on OpenRouter rose from 0.51 trillion to 7.3 trillion tokens in six months, a 14x increase, versus 2.8x growth for human usage, according to OpenRouter data compiled by Peter Walker and circulated in a16z's Charts of the Week in August 2026. February 6, 2026 was likely the last day humans consumed more tokens than agents, and OpenAI's Enterprise Signals report released August 13, 2026 found Codex accounted for 64% of combined Codex and ChatGPT enterprise output tokens in June 2026. Goldman Sachs Research projects consumer and enterprise agents will consume 120 quadrillion tokens a month by 2030, 24 times the 2026 level.

by read3 min views1 publishedSep 7, 2026
The Three Waves of AI Consumption
Image: Tomtunguz (auto-discovered)

On February 6, 2026, agents on OpenRouter consumed more tokens than humans did. They have not given the lead back<sup>1</sup>. AI consumption is not a smooth curve anyone can extrapolate : it arrives in three waves, each orders of magnitude larger than the last. The second one has already broken. The third is on the horizon.

Wave one is chat : a person asks, the model answers once, done. Roughly 1m tokens per active user per day, by our estimate.

But chat is already the minority of enterprise output : by June 2026, chat accounted for a little more than a third of enterprise use, & Codex for the other 64%<sup>2</sup>.

Wave two is a single agent vibe coding a family weekend sports scheduling app to coordinate carpools, or researching trends in the ten-year bond & the macroeconomic implications. Roughly 100m to 200m tokens a day, again an estimate. The token intensity gap between waves is two zeros of power<sup>3</sup>.

Wave three stacks a meta-harness on top : one AI dispatching many agents in parallel, each spawning its own tool calls & sub-agents. Another order of magnitude, & the numbers hit the billions.

Consumption does not grow because generation gets quicker. It grows because parallelization compounds.

Here’s another way of building a meta-harness. Amass a spreadsheet not of numbers in each cell, but of AI answers. Imagine a list of YCombinator startups 200 long with 30 columns each : summary of founder backgrounds, history of the company, value proposition. Each column represents tens of AI tool calls.

The run costs tens of millions of tokens before anyone notices.

That is the shape of the shift : agent volume on OpenRouter rose from 0.51t to 7.3t tokens in six months, a 14x increase, while human volume managed 2.8x<sup>1</sup>.

The demand implied by these waves is already showing up at the macro level. Goldman Sachs projects consumer & enterprise agents will consume 120 quadrillion tokens a month by 2030, 24 times the 2026 level<sup>4</sup>.

The waves do not replace each other, they stack. Chat keeps growing, agents grow faster, meta-harnesses will dwarf both. Anyone sizing compute off a smooth extrapolation is planning for the wrong curve.

OpenRouter data compiled by Peter Walker, circulated in a16z’s Charts of the Week, August 2026 : agent token consumption rose from 0.51t to 7.3t tokens, a 14x increase, against 2.8x for human usage; 6 February 2026 was likely the last day humans consumed more tokens than agents. Agents consume nearly5x more tokens per task than humans on the same models. Two caveats : the large majority of agent token burn is cached prompts billed at lower rates, & OpenRouter’s mix skews toward open-weight models.↩︎↩︎ 2. OpenAI : Enterprise Signals , report released 13 August 2026 : Codex accounted for 64% of combined Codex & ChatGPT output tokens among enterprise customers in June 2026, leaving ChatGPT at 36%. The report tracks January to June 2026; frontier firms used 8.3x more tokens than typical companies by June, up from 2.6x in January.↩︎ 3. Chat runs about 1m tokens per active user per day against 100m to 200m for a single agent, roughly two orders of magnitude. Measured per task rather than per day the gap is wider still : Yu et al., How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks finds agentic tasks consume roughly 1,000x more tokens than code reasoning & code chat. Input tokens drive the total, because an agent re-reads its own context on every step.↩︎ 4. Goldman Sachs Research : AI Agents Forecast to Boost Tech Cash Flow as Usage Soars , May 2026 : token consumption rises 24x to 120 quadrillion tokens a month by 2030 across consumer & enterprise agents.↩︎

── more in #ai-agents 4 stories · sorted by recency
── more on @openrouter 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-three-waves-of-a…] indexed:0 read:3min 2026-09-07 ·