cd /news/artificial-intelligence/tokenmaxxing-is-out-valuemaxxing-is-… · home topics artificial-intelligence article
[ARTICLE · art-101286] src=fastcompany.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Tokenmaxxing is out, valuemaxxing is in

Tesla capped employee AI spending at $200 per week after six months of ranking engineers on internal AI leaderboards by token usage, joining Uber, Meta, Amazon, and Walmart in reversing course on unlimited AI use. The shift from 'tokenmaxxing' to 'valuemaxxing' comes as flat-rate plans that sold tokens below cost expire, with power users able to burn through $14,000 in tokens on a $200-per-month plan, exposing a 4,500-times price difference between the cheapest and priciest AI models.

read5 min views1 publishedAug 18, 2026

It may be game over for gamified token consumption. Tesla spent six months ranking its engineers on internal AI leaderboards by token usage, then thought better of it and capped employee AI spending at $200 per week. This should sound familiar. Uber, Meta, Amazon, Walmart, all reversed course in the same direction.

How did we get from all-you-can-prompt to token-pinching? Think about it this way.

You land at an airport, call an Uber, and feel glad you never had to park a car. Sure, it costs more, but the convenience is worth it.

For your daily commute, you rely on your own car and deal with the inconveniences. There’s a simple, almost instinctual logic at play. You pay more when it’s worth it, and you save when it’s not. It’s not quantum computing. And yet, somewhere along the way, the simple economic fact of scarcity broke down and mutated into the illusion of superabundance.

Tokens are the fuel AI systems burn, and there is a 4,500-times price difference from the cheapest to the priciest AI model. Picture pulling up to a gas station and finding two pumps. Regular is under $4 a gallon. Premium reads $17,000 per gallon.

Life rarely presents us with such easy decisions to just forget about premium! Except there’s a confounding variable where AI is concerned.

It turns out you must use both fuel sources, the nozzle defaults to premium, and your dashboard won’t reveal which pump you’re on. You might be filling the premium tank by asking an LLM a question as silly as if you should wear cargo shorts today. I should know because I’ve done it. Your teams have done it too.

All that tokenmaxxing was harmless when the cost seemed minimal. Now that the finance team has started counting, they’ve had to push their eyeballs back in their heads and ask people like me the question everyone skipped two years ago: what return are we getting for all this?

The question never came to the surface in the past because the cost was buried. Flat-rate plans sold tokens below cost, the labs covered the difference at the pump, and now that subsidy is running out. Push the most popular chatbot’s $200-a-month plan to its limit,, and a power user can burn through $14,000 in tokens at list price. The other $13,800 was the subsidy. And here’s the crazy part. When you hit your limit mid-task, you go get a coffee and wait. An agent running independently at 3 a.m. can’t do that, so teams hand it unlimited capacity and walk away. Nobody’s watching the pump when it’s pumping the fastest. Weekly caps, rate limits, locking out users mid-session, neutering subscription tiers—these are the subsidy being pulled back in public view. It’s all prelude to the shift from tokenmaxxing to valuemaxxing.

Valuemaxxing is what it sounds like. You spend where the model earns its return on investment and the rest gets routed to something cheaper.

The goal is to match what you pay to what you get back—one task at a time. There’s nothing magical about valuemaxxing. It’s basically using the right tool for the right job.

Here are the three steps for valuemaxxing.

**1. Put a gauge on your dash. **You can only spend by value if you can see what you’re spending. That gauge has a name: FinOps, financial operations, and it is a sibling to DevOps. DevOps ships the work while FinOps never loses sight of the costs.

**2. Stop buying premium for your daily commute. **It’s the culinary equivalent of putting foie gras on a fast-food burger: wrong tool, wrong job. Send all your routine, high-volume work—mundane tasks such as asking whether you should wear cargo shorts today—to something cheaper: open models or the labs’ own budget tiers. Reserve the frontier power for the jobs that require agentic, multi-step capabilities, especially since those can burn 1000-times the tokens of a single chat prompt.

**3. Know when it’s time to stop calling Uber and buy your own car. **Past a certain monthly volume, the per-token cost of hardware you own can fall well below what you’d pay the biggest AI providers, and at real scale it starts paying for itself. Right model, right location. The pump is one decision, and whose engine you’re renting is the other. There’s a time for Uber and a time for your own car.

Get all three right, and AI stops being an expense you tolerate and becomes an investment you control.

Make those three moves and you won’t need to worry about the next “-maxxing” trend. Per-token prices may keep falling, but whether your bill will follow is a different question. Cheaper tokens could drive more consumption, and then the efficiency gets spent as fast as it comes. What doesn’t change is that most of your work is built for the commute. It’s routine, high-volume, and it runs great at the cheaper pump.

For the big lifts, such as modernizing a legacy code base or a fleet of customer-facing chatbots running around the clock, the trip is worth the premium upgrade. The agentic intervention earns the frontier prices—and the keyword is earns. The routine work is about 80% and the big lifts about 20%. That split allows you to cut your bill without losing anything you can measure.

So no, valuemaxxing isn’t token minimizing. The winning strategy has never been to use less AI. Go bonkers. Use as much as you need. Burn 6 million tokens on a Tuesday morning, but if, and only if, the work is worthy of the spend.

Check the gauge before you fill up and know which pump you’re reaching for. We’ve spent two years forgetting we already knew the answer: A place for every token, and every token in its place.

Juan Orlandini is chief technology officer of North America for Insight Enterprises.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @tesla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tokenmaxxing-is-out-…] indexed:0 read:5min 2026-08-18 ·