cd /news/ai-agents/the-agentic-code-review-loop · home › topics › ai-agents › article
[ARTICLE · art-148239] src=robinwieruch.de ↗ pub= topic=ai-agents verified=true sentiment=· neutral

The Agentic Code Review Loop

A freelance AI engineer introduced a three-week agentic code review loop to a product team in which every developer maintains a personal AI review skill that fires at teammates' pull requests and posts findings as inline comments, while the author's own agent triages each finding by fixing, declining, or refuting it. The setup deliberately avoids a single shared repository reviewer to preserve diversity of findings, and a weekly scorecard ranks which developers' review skills perform best. The author reports the remaining weak spot is that someone must still manually track open pull requests, which early team members are automating with a top-level agent that polls every 15 minutes and launches a subagent per unreviewed pull request.

read12 min views1 publishedOct 9, 2026

A few months ago I wrote about how we let AI agents review code against our documented patterns, because the review bottleneck had moved from writing code to reading it. That worked, and then it created the next problem: once every developer has a review agent, who reviews the reviewers?

Read MoreAgentic Code Review: Pattern Matching for AI

As a freelance AI engineer, I introduced a review process to a product team over the last three weeks. Every developer owns a personal AI review skill. They fire it at each other’s pull requests and walk away. The author’s agent answers. And at the end of the week, a scorecard tells us whose review skill is actually any good. This post walks through the loop, the numbers from the first three weeks, and where it still has weak spots.

The obvious way to do AI review is one shared reviewer for the whole repository. One prompt, one configuration, maintained by whoever cares most. We deliberately did not do that.

Instead, every developer keeps a review skill on their own machine: a set of instructions their coding agent follows when asked to review a pull request. One developer’s skill hunts for duplicated code, another one reads the API contract first, a third one runs an adversarial pass. Nobody is asked to align. The reasoning is the same as for human review: five people who think alike find the same five things. A single shared reviewer also overfits. It gets tuned against the findings the team already knows to look for, and it stops surprising anyone.

The second rule is fire and forget. When a pull request opens, every teammate is invited to point their review skill at it. The agent posts its findings as inline comments, and the developer who started it does not look at them again. That sounds careless, and for a human review it would be. But it is the only way the volume works. A reviewer who has to read their own agent’s review first is back to being the bottleneck.

Each comment says which model wrote it, so nobody mistakes it for a colleague’s opinion:

[Model Name]:
> [SUGGESTION] `user-list.tsx:56` The empty check is repeated
> in every list component. The base list already knows whether
> it is , so it can make this decision itself.

Fire and forget still leaves one chore: someone has to fire. Right now every developer keeps track of the open pull requests themselves, which ones they already reviewed and which ones are ready for another round. The first developers on the team are automating that with one top-level agent running in a loop. Every 15 minutes or so it checks for pull requests it has not reviewed yet, and for each new one it launches a subagent that runs the whole review.

The alternative would be to host everyone’s review skill in the cloud, so that reviews arrive on their own the moment a pull request opens. I do not think that works for us. A review skill is only half of a reviewer. The other half is the harness around it, the developer’s own agent setup, and that does not move to the cloud with a prompt file. It would also put every skill in one shared place, which works against the diversity we wanted, and it would move the token cost away from the developer who chose to run the review.

Fire and forget moves the work to the author. A pull request can collect twenty, thirty, sometimes fifty findings from different agents, and no human reads fifty findings carefully.

So the author answers with an agent as well. It takes each finding, checks it against the actual code, and decides: fix it, decline it because it is not worth the change, or refute it because the finding is wrong. Then it leaves a reply in the thread, again marked as written by a model. A fix names the commit. A refutation brings the evidence.

> Non-issue: the field is required by the schema, so an empty
> value is rejected before this branch can run.

Written by Model Name

The answering agent has the harder job of the two, and it cannot be a shallow one. A single agent that walks through a hundred findings in one session fills up its context window long before the end, and from then on it does not judge findings anymore, it waves them through or waves them off. What works better is handing each finding, or a small batch of related ones, to its own subagent that starts fresh, reads the code, and comes back with a verdict. I can recommend that to the team, but I cannot prescribe it. How a developer’s agent answers is as much their own setup as how it reviews.

If the first round produced many findings, a second round is fair game after the fixes landed. One rule turned out to matter a lot here: a later round has to read the earlier threads first. Without that, one agent recommends moving to approach A, the author moves, and the next agent recommends moving back to B. Two agents without shared context will happily argue a pull request in circles.

  1. A pull request opensThe author asks the team for reviews, as always.
  2. Every teammate fires their own review skillFire and forget: each agent posts its findings as inline comments, and the developer who started it moves on.skill Askill Bskill C
  3. The author's agent answers every findingIt checks each finding against the code, fixes what holds up, and leaves a reply in the thread either way.fixeddeclinedrefuted
  4. Another round, if the first one found a lotA later round reads the earlier threads first, so it does not argue the code back to where it started.
  5. A human reviews lastIs this the right change, was it asked for, does the design fit? The pattern matching is already done.

None of this replaces the human review. It changes what the human review is for.

After the agents are done, a developer still reviews the pull request and puts their name on it. They are free to use their agent to explore the change, ask questions, check a hunch. But it is not fire and forget anymore. The inline comments are theirs, and they stand behind them.

What they no longer have to do is the pattern matching. Naming conventions, a missing test, a duplicated helper, an unhandled error path: the agents found those. The human question is a different one. Is this the right change? Was this feature asked for? Does the design fit where the product is going? Those are questions an agent answers confidently and badly.

Read MoreYour AI Output Is Someone Else's Input

Here is the part that made the whole setup interesting to me. Every finding has an author (a review skill) and, ideally, a verdict (the reply). That is a dataset.

After three weeks I sat down with an agent and scraped it: every inline review comment on every pull request in that period, grouped by the developer whose skill posted it. A second pass read the replies under each finding and classified the outcome as valid, declined, or refuted. The result was 789 findings from five review skills. Of the findings that got a clear answer, about three quarters were accepted and 8% were refuted.

The averages are not the interesting part. The spread is. The totals above are real. For the breakdown below I made up the names and adjusted the numbers per skill, so that no teammate is recognizable and every outcome a scorecard like this can produce shows up once:

Each of the five skills tells a different story:

  • Mara’s skill is never wrong, and it almost never says anything. A typical review posts one or two findings, and on pull requests that other agents also reviewed, it misses most of what the others catch.
  • Jonas’s skill sits where everyone wants to be: reasonably deep and right most of the time.
  • Lena’s skill finds far more than everyone else, but a quarter of what it posts gets declined as not worth the change. It buries the author in nits.
  • Tariq’s skill is thorough and often wrong: a quarter of its answered findings were refuted.
  • Felix’s skill is both shallow and imprecise.

Precision alone would have put the wrong skill on top. The skill with a perfect score is the one I would improve first, because a reviewer who only speaks when certain is leaving most of the review undone. This is why the scorecard needs both axes: how often a skill is right, and how much it finds when several skills look at the same change.

It also gives every developer something concrete to do. The shallow skill needs to look for more. The one that is often wrong needs a verification step before it posts. The deepest one needs to stop posting every nit. Nobody has to agree on a shared prompt for that. Everyone tunes their own.

This is the kind of verdict the scorecard hands to each developer:

Next week’s scorecard shows whether the change worked. That is the whole idea: a developer does not have to guess anymore whether a tweak to their review skill made it better.

The scorecard has a weak spot, and it showed up immediately.

Only about half of the findings were ever answered. Many of the rest were clearly handled: the commented lines changed afterwards and the thread was resolved. But a changed line is not a verdict. Maybe the finding was right. Maybe that code was rewritten for a different reason. To know, an agent would have to read the follow-up commits for every unanswered finding, and that burns far more tokens than reading a one-line reply.

The reply rate differed wildly per author, from almost every finding answered to almost none. That skews the score. A review skill whose findings mostly land on a careful responder gets refuted more often than one whose findings land on someone who fixes quietly.

So the one rule this process really needs is small: every AI finding gets a reply. It does not matter who writes it, the developer or their agent, and it does not matter how it is worded. We decided against a fixed vocabulary, because aligning every developer’s replying agent is exactly the kind of alignment the setup tries to avoid. A model can read any reply and tell whether it says fixed, declined, or wrong. It cannot read a reply that is not there.

I would not trust this scorecard blindly, for two reasons.

The first one is Goodhart’s law. The moment a review skill is rated by its acceptance rate, the cheapest way to improve the rating is to post fewer, safer findings. The chart above already shows what that looks like: two of the five skills rarely post more than two findings per review. Where we want to end up is the opposite, exhaustive reviews instead of shallow ones. The depth axis is the counterweight, and it only works on pull requests that several skills reviewed.

The second one is the judge. Right now the verdict comes from the author’s agent, and the author is the party that saves work when a finding gets declined. The score does not measure whether a finding was correct. It measures whether two agents agreed. What is missing is a judge outside the loop: something that runs once when a pull request merges, belongs to neither the reviewer nor the author, and gets spot-checked by a human every now and then. We do not have that yet.

There is a third number I would like to have and cannot get from the threads at all: what did the human review, or production, catch that no agent flagged? That is the only real measure of whether the loop misses things.

If the measuring holds up, the direction is a self-improving loop. Review skills get scored every week, developers tune them, precision and depth go up together, and the human review gets shorter because there is less left to find.

One step further sits a classifier at the end of a pull request that says how confident it is that nothing is left. A small change with few findings and a high score could merge on its own. A large one could earn its way there through review rounds. Architecture and security changes would stay with a human no matter what the score says.

I want to be careful here, because that is an outline and not a plan. A model’s confidence is not calibrated just because it is a number. Before any score merges anything, I would record it silently for a while and compare it with what the human review still found. And there is a quieter risk on the way: when five agents have already approved a pull request, the human review turns into a rubber stamp long before anyone decides that it should.

What convinced me of this setup is that it became measurable. Three weeks ago, “my review agent is pretty good” was a feeling every developer on the team had about their own setup. Now it is a number with a known weakness, next to four other numbers, and the weakest part of the whole loop turned out to be a missing one-line reply. That is a much better problem to have than the one we started with.

── more in #ai-agents 4 stories · sorted by recency
── more on @user-list.tsx 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-agentic-code-rev…] indexed:0 read:12min 2026-10-09 · —