# Org Leaders: Is Your AI Policy Attracting the Best Talent?

> Source: <https://blog.kilo.ai/p/org-leaders-is-your-ai-policy-attracting>
> Published: 2026-07-29 00:48:19+00:00

# Org Leaders: Is Your AI Policy Attracting the Best Talent?

### Developers don't want to work with locked-in tooling.

Jensen Huang said recently that he expects an employee earning $500,000 to consume $250,000 worth of tokens, and he called token allowances “one of the recruiting tools in Silicon Valley.”

Around the same time, Andrej Karpathy admitted he gets nervous when he *doesn’t* burn through his AI budget, and that when he hits a quota on one tool, he simply switches to another.

Now picture the engineer you’re trying to hire in that world. One offer comes from a company where AI usage is firmly encouraged, because they’ve built the governance, model flexibility, and cost visibility to make freedom affordable. The other comes from a company where a $1,500 monthly cap kicks in around week two and over-limit requests need a manager’s sign-off.

The salary and title are identical, so where does that engineer go?

This isn’t hypothetical anymore. Uber burned its annual AI coding budget in four months and responded with per-employee caps. Microsoft’s Experiences and Devices division gave thousands of engineers Claude Code, watched token billing consume the annual budget ahead of schedule, and mandated a migration to a cheaper tool. Gartner is now forecasting that AI coding token costs will overtake the average developer’s salary by 2028, and their peer data shows 6% of organizations already spend more than $2,000 per developer per month.

To be clear, the cost concern is legitimate. Nobody should be handing out blank checks, and the “tokenmaxxing” era of ranking employees by consumption deserved to die. But the response most organizations reached for, hard caps on a single premium provider, solves the spend problem by creating a talent problem.

The question worth asking isn’t “how do we limit usage.” It’s “did we build the tooling and cost optimizations that let us give developers access to everything they want to use?” Those are very different engineering problems, and only one of them shows up in your offer letters.

**What developers actually do when nobody constrains them**

At Kilo, we sit in an unusual position. We’re a model-agnostic platform, so we don’t force developers into one provider, one model, or one workflow. That means our telemetry, drawn from roughly 3 million developers and more than 40 trillion tokens, shows what engineers do when they’re free to choose. Most usage data in this industry comes from the labs themselves, and a lab can only see its own traffic. We see the whole field.

The top-line finding is that developers do not want one model. Over the past year, the share of developers using more than two providers monthly grew to 68%, and that population is growing about 6% month over month. More striking: 45% of those developers switched models within the same hour.

That hourly switching matters because it exposes a mismatch in how organizations think about AI. Procurement treats model choice as an annual or quarterly decision made by leadership. The telemetry says the real decision happens constantly, at the level of individual tasks, made by individual engineers based on whatever is in front of them right now.

And here’s the part that should reassure your finance team: it’s not aimless. About 38% of switching is developers trading cost against quality, moving between paid and free models depending on the task. Another third is paid-versus-paid, engineers checking whether the expensive model actually earns its price on their specific work. The rest is people rotating through free models to find the best result for nothing.

Read that at any altitude and it looks like engineers running cost-benefit analysis by hand, one task at a time. Given the freedom to choose, developers don’t maximize spend. They optimize it. Your job as a leader is to give them the tooling that makes that optimization automatic instead of manual.

**Your most experienced engineers are the most multi-model**

When we bucket developers by usage volume and count how many providers they touch, the relationship is nearly a straight line. Light users, under 100 requests a month, average about 1.5 providers. Heavy users, above 10,000 requests a month, average more than 7, and 92% of them use multiple providers.

The tempting read is that power users are outliers. I’d argue it’s backwards. These are the people who’ve logged the most hours, hit the most edge cases, and built the most accurate sense of what each model is actually good for. They’re the leading indicator of where everyone else lands. Single-provider usage isn’t a stable state that developers settle into. It’s a starting point that wears off with experience.

So when an organization standardizes on one provider or slaps a hard cap on usage, it’s effectively locking in the behavior of its newest users and overriding the judgment of the people who’ve used these tools the most. Those are precisely the engineers with the most options on the job market.

If you think managed enterprise environments are immune to this, our data on managed org seats says otherwise. Multi-provider usage on team subscriptions went from 42% to 71% in a single quarter. Same curve as individual developers, just running about two quarters behind.

The sentiment data lines up. The Pragmatic Engineer’s 2026 survey of over 900 engineers found roughly 30% hit usage limits monthly, and respondents described running out of tokens mid-task as one of the most disruptive things that happens to them, precisely because it breaks flow state.

Meanwhile, Stack Overflow’s 2025 survey of 49,000+ developers found the top driver of job satisfaction isn’t compensation. It’s autonomy and trust. A usage cap is a very legible signal about both.

**The cost problem is real, but it’s a routing problem, not a usage problem**

The reason so many organizations got burned is that they standardized on a single premium provider and then discovered what a monoculture costs at scale. Our data shows exactly how much money is sitting in the gap between the benchmark leaderboard and what work actually requires.

At the time we pulled this data, GPT-5.5 sat at number one on our completion benchmark at 74.2%, but ranked 19th by actual token volume. Claude Opus 4.7 was second on the benchmark and 20th by usage. The most-used model on the platform in a given week was a free model most engineering leaders have never evaluated. Part of the reason is cost: the top benchmark model runs about $72.63 per attempt versus a floor of $20.65 among the rest of the top 10. That’s a 3.5x price difference for roughly 20 points of completion rate, and whether those points are worth it depends entirely on the task, which no leaderboard can tell you.

Zoom out to the full year and the story holds. MiniMax, not Anthropic, Google, or OpenAI, was the most-used lab on Kilo, about 21.5% ahead of Anthropic. Anthropic’s share of requests dropped from 11% to under 7% over an eight-week window while their user count roughly tripled, because the whole field grew faster than any one lab. On any given day, the top 10 most-used models spanned 6 different providers, and 37,000 developers tried a brand-new provider in the most recent week alone.

This is a fragmenting market, not a consolidating one, and there’s a structural reason. Search and cloud consolidated because of network effects and real switching costs. The model layer has neither. One model doesn’t improve because a competitor uses it, and switching is close to free when the tooling supports it. When switching is cheap and lock-in doesn’t exist, the economics push toward more options.

For an org leader, that fragmentation is the whole opportunity. Most of the work flowing through your engineering org does not need the $72 model, and your developers already know it. The organizations with runaway bills aren’t the ones using AI the most. They’re the ones paying frontier prices for boilerplate.

**What the tooling actually looks like**

Closing the gap between “capped” and “unconstrained” is an infrastructure decision, and the pieces are more concrete than most leaders assume. This is the stack we built Kilo around, but the principles apply regardless of what you’re evaluating.

**Automatic routing.** The cost-quality tradeoff our developers make by hand is exactly the kind of decision that should be automated. Kilo’s Auto Efficient does session-aware model routing, matching each task to the cheapest model that can actually handle it, so the optimization happens without anyone breaking flow to think about it.

**Real model freedom.** With 500+ models available and switching at any point in a task, engineers can route routine work to free and low-cost models and reserve premium inference for the problems that justify it. A monoculture can’t do this by definition, and the free tier of the model market is far better than most procurement processes give it credit for. The most-used model on our platform costs nothing.

**Transparent, at-cost pricing.** Flat seats hide the real economics. Per-token pricing at provider cost means your finance team sees exactly where spend goes, which makes “unlimited” a measurable engineering target instead of a leap of faith.

**Centralized visibility without centralized restriction.** Pooled credits and unified billing prevent shadow IT sprawl, and an AI ROI dashboard tells you what adoption and velocity you’re actually getting for the spend. The FinOps Foundation found 98% of practitioners now actively manage AI spend, up from 63% a year earlier, so the visibility layer is table stakes. The differentiator is whether that visibility feeds optimization or just rationing.

The point of all of it is a specific outcome: developers who never think about limits, and a finance team that never gets surprised. Those two things are only in tension if your stack can’t route.

**The recruiting question**

Put the pieces together and the orgs that win the next hiring cycle look like this: they can tell a candidate that AI usage is simply not something they’ll think about, not because the budget is infinite, but because the blended cost per task has been engineered low enough that caps stopped making sense.

That posture compounds. Your senior engineers keep their judgment intact instead of routing around policy. Your juniors ramp faster because experimentation costs nothing. Your recruiting pitch includes a line your capped competitors can’t match, and your finance team gets better data than the flat-seat orgs, not worse.

Huang is right that tokens are becoming a recruiting tool, but the orgs that come out ahead won’t be the ones that spend the most. They’ll be the ones that did the unglamorous tooling work, routing, model freedom, and cost visibility, so that offering developers everything became affordable. Our telemetry says the most experienced half of the developer population is already living in that multi-model world, and your managed org is trailing them by about two quarters on the same path. So if you’re an org leader, it’s worth asking the question the way a candidate will: does your AI policy read like infrastructure, or like rationing?
