cd /news/ai-tools/dont-measure-ai-code-percentage-on-i… · home › topics › ai-tools › article
[ARTICLE · art-146701] src=newsletter.getdx.com ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Don’t measure AI Code Percentage on its own

AI-authored code reached 52% of all merged code by Q2 2026, up from 19% in Q3 2025, according to DX research across more than 500 companies where weekly active AI usage now exceeds 95%. DX researcher Brian Houck argues AI Code Percentage should be used as a research variable to segment teams rather than as a KPI, since teams can inflate it by routing trivial ten-line changes through agents. Houck said the metric "is more useful for explaining outcomes than as an outcome itself.

by read7 min views3 publishedOct 7, 2026
Don’t measure AI Code Percentage on its own
Image: Newsletter (auto-discovered)

Welcome to the latest issue of Engineering Enablement, a weekly newsletter sharing research and perspectives on developer productivity.

Hi everyone, I’m Scott, the new Managing Editor at DX.

Each week I’ll share what I am seeing from across the industry, with a particular focus on AI enablement, metrics, and software team dynamics. You’ll still regularly hear from Brian Houck, Justin Reock, Eirini Kalliamvakou, Gratiana Fu, and our extended research community.

This week, I’m going to hand things over to Brian, who’s been thinking about the popular AI code percentage metric, and why, in a world where almost every developer has adopted AI coding tools, it needs to be applied as a lens and not a KPI.

Brian: The question of whether developers are using AI, and whether it is writing a lot of their code, is settled.

Across more than 500 companies, weekly active AI usage now exceeds 95%. And the percentage of code being generated by AI keeps climbing. In Q3 of 2025, an average of 19% of all merged code was AI-authored. By Q2 of 2026, that figure had reached 52%.

When AI was new, the percentage of AI-authored code was a useful proxy for adoption. Moving from 10% to 30% meant developers were working differently than they had been, and it was fair to read that as progress. However, adoption is no longer an open question, so the metric has lost the role it was originally intended for.

That doesn’t mean AI Code Percentage (AICP) should come out of our measurement systems. It means being clear about its role in them.

A useful research variable can be a bad target #

This distinction matters to me as a researcher. I want to know how much code is being generated by AI.

AICP lets us split developers or teams into cohorts and ask much more interesting questions. As AI-authored code increases, what happens to:

  • PR throughput and cycle time?
  • Change failure rate and other quality measures?
  • Code review wait time?
  • Time spent on new capabilities versus maintenance?
  • Developer experience?
  • Developers’ confidence in the changes they’re shipping?

Most of those map onto dimensions the Core 4 already tracks: Speed, Effectiveness, Quality, Impact. AICP isn’t one of them. It’s the variable you segment by in order to answer why those dimensions may be changing, which makes it an input to better understanding engineering performance, rather than a measure of engineering performance itself.

In research terms, AICP is often more interesting as an independent variable than a dependent variable. For engineering leaders, I’d put it more simply:

AI Code Percentage is more useful for explaining outcomes than as an outcome itself.

Being a good diagnostic doesn’t make it a good target. The test I apply to any metric someone wants to use as a goal is whether gaming it still produces the outcome you wanted. Time-to-First-PR passes this test. Push a new hire to open a trivial PR in their first week and you’ve still forced the onboarding environment to work, which was the thing you cared about. Gaming it and doing it properly are hard to tell apart, and that’s what makes it a good target.

AICP fails badly. Tell a team its AICP needs to reach 60% and they can deliver that within a quarter by routing work through agents that would have been a ten-line human change. Benchmarking also falls apart. Suppose the measurement was perfect and you knew your organization sat at 40%, while a peer sat at 60%. You’d have learned that two companies produce code differently. You still wouldn’t know which one gets more value from AI, and the gap on its own gives you no reason to move in either direction.

Activity metrics aren’t the problem #

One of the central ideas behind the SPACE framework was that developer productivity cannot be reduced to a single measure, and that activity should not be confused with productivity.

Activity data isn’t useless. Quite the opposite. Commits, pull requests, deployments and other observable behaviors can tell us a great deal about how a software engineering system works.

Though on the other side of the argument is lines-of-code, which has earned its reputation as a meaningless activity metric. Producing twice as much code doesn’t mean you’ve produced twice as much value, or even that you’ve done twice as much work. Often the best solution is the one that needs the least code.

AICP sidesteps that trap. It isn’t a count, it’s a rate. Since it is a proportion, more code doesn’t inherently mean a higher score, and a larger team doesn’t automatically outperform a smaller one. The question shifts from “How much code are we producing?” to “How much of our code production involves AI?”

AICP does share one limitation with lines of code: a higher number doesn’t necessarily mean a better outcome. It tells us how much of our code production involves AI, not how effectively we’re using it.

The same AI Code Percentage can describe very different ways of working #

This problem gets more complicated as we move from autocomplete toward agents.

Imagine two teams that both report 60% AI-authored code. On the first team, developers use AI primarily as sophisticated autocomplete. They decide what to build, decompose the problem, write most of the implementation logic, inspect suggestions as they’re produced, and retain the useful ones. On the second team, developers delegate entire tasks. They describe the desired outcome, an agent explores the repository, implements a solution, runs tests and submits a pull request.

Both teams might produce the same AICP, but the role AI plays in their work is radically different.

This is why autonomy is becoming a more interesting dimension of AI adoption. Knowing that AI produced a piece of code doesn’t tell us how much of the surrounding task AI performed, how much human guidance it required, or where human judgment entered the process.

Generated isn’t the same as shipped #

Suppose an AI agent generates 1,000 lines of code. A developer accepts 800 of them, rewrites 300 during review, removes another 200 before merge, and six weeks later half of what’s left has been replaced.

How much AI-authored code was there? The answer depends heavily on when you measure it.

Attributing code to AI is inherently messy. AI generates code that humans modify. Humans move and refactor AI-generated code. Suggestions get partially accepted. Developers work across multiple tools and interfaces. Agents behave differently from IDE autocomplete.

This is why I think survivorship is an increasingly important companion to AI Code Percentage. A useful way to think about AI output is as a funnel:

Recent research is beginning to look at exactly this question. A 2026 study of more than 200,000 code units across 201 open-source projects examined the subsequent fate of agent-authored code rather than simply measuring how much was generated.

The researchers found meaningful differences in how AI and human-authored code evolved after merge. More interestingly, they found significant differences in survivorship between different AI tools. The survival rate of code from Claude was nearly twice that of code from Devin.

How I’d actually use it #

At DX, we think AICP should be treated as a lens, not a KPI. Read it alongside throughput, cycle time, quality, time allocation, and developer experience. If AICP rises, look for whether teams are delivering faster, maintaining quality, and spending less time on work they find tedious. If those metrics don’t improve, perhaps AI is exposing new bottlenecks. The patterns won’t establish cause and effect on their own, but they tell you where to investigate.

Review it quarterly, not weekly. This is the part I see missed most often. AICP describes a structural property of how a team works, and structural properties move slowly. A metric earns a weekly slot when you expect it to respond to something you did last sprint. AICP mostly won’t. Watching it closely generates noise, and it creates a temptation to force the line to move upwards.

As agentic development matures, we’ll need measures AICP was never designed to provide: how much autonomy agents are given, the quality of the context and requirements they receive, how reliably they accomplish what the developer actually wanted, and what happens to their output after it merges.

AICP has a place in that system. It just shouldn’t be the point of it.

Don’t ask, “How do we increase our AI Code Percentage?” Ask, “What changes when our AI Code Percentage increases?”

A lens only works if you’re looking through it at something else.

That’s it for this week. Thanks for reading.

-Scott

── more in #ai-tools 4 stories · sorted by recency
── more on @dx 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dont-measure-ai-code…] indexed:0 read:7min 2026-10-07 · —