cd /news/ai-agents/four-horsemen-of-agent-prs-and-how-t… · home › topics › ai-agents › article
[ARTICLE · art-144296] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↓ negative

Four Horsemen of Agent PRs and How to Stop Them

A forensic study of 33,000 agent-authored pull requests found agents achieve an 83.77% acceptance rate versus 91.01% for humans, but fail in fundamentally different ways: syntactically correct code that violates contracts, breaks cross-service dependencies, and introduces architectural regressions that pass every test. The analysis identifies four recurring failure modes, dubbed the "Four Horsemen," with context collapse — locally correct but globally wrong code — cited as the most pervasive, and notes that 23% of rejected agent PRs were duplicates of work already underway.

by read9 min views1 publishedOct 3, 2026

Agent-generated pull requests grew from under 1% of GitHub PRs to 27.6% in just fourteen months. Anthropic's 2026 Agentic Coding Trends Report puts the number higher: 41% of all new code is now AI-generated, with tools like Claude Code growing 6x in workplace adoption in under a year. The PR flood is here. The review infrastructure is not.

📖 Read the full version with charts and embedded sources on AgentConn → The problem is not that agents write bad code. It is that agent-authored PRs fail in ways human-authored PRs never did, and our entire review process was designed for a world where a human could explain their reasoning when asked. A forensic study of 33,000 agent-authored PRs found that agents achieve an 83.77% acceptance rate versus 91.01% for humans, but the gap hides the real story: the types of failures are fundamentally different. Agents do not make typos or forget semicolons. They produce syntactically correct, compilable code that violates contracts, breaks cross-service dependencies, and introduces architectural regressions that pass every test in the suite.

After months of tracking this data across academic research, industry reports, and practitioner experiences, we see four recurring failure modes that account for the overwhelming majority of agent PR disasters. We call them the Four Horsemen.

The agent knows the file. It does not know the system.

The first and most pervasive failure mode is context collapse: the agent generates code that is locally correct but globally wrong. It fixes the function but breaks the contract. It refactors the module but violates the architectural boundary. It adds the feature but duplicates logic that already exists three directories away.

FeatBit's analysis of the 2026 productivity paradox identified four types of context loss in AI-generated pull requests: requirement context (what the business actually needs), codebase context (how the system fits together), review context (what previous reviewers flagged), and organizational context (team conventions, deployment constraints, implicit rules that live in people's heads).

The MSR 2026 empirical study found that 23% of rejected agent PRs were duplicates, with agents frequently submitting PRs for issues already being addressed by another contributor. The agent had no idea someone else was working on the same problem. It had no concept of "someone else."

View original post on Stack Overflow Blog → The Stack Overflow engineering blog quantified the damage: AI-generated code produces 1.7x as many bugs as human code, with logic and correctness errors running 1.75x higher and security findings 1.57x more common. But here is the insight most teams miss: these are not bugs of incompetence. They are bugs of context. The agent wrote correct code for the wrong problem because it could not see the full system.

⚠️ The context collapse trap: Agents excel at small, well-defined changes. The MSR study found that about 28% of agent PRs merge almost instantly. The danger zone is anything that touches multiple services, crosses architectural boundaries, or requires understanding of business rules that live outside the codebase.

How teams are surviving it:

domain/ may import from infrastructure/" catches the boundary violations agents love to introduce. It looks right. It reviews right. It breaks in production.

View original report on New Relic → This is the most insidious horseman, and the New Relic 2026 State of AI Coding report gave it a number: 94% of engineering leaders rate AI-generated code as higher quality than human code at the time of review. The code reads well. The variable names are descriptive. The comments are helpful. The structure is clean.

Then it hits production. 78% of those same respondents report more incidents once deployed. 82% experienced at least one production failure tied to AI-generated code in the past six months. 74% say at least a quarter of AI code needs significant rework within twelve months.

The paradox is structural: agents are optimized to produce code that looks correct to a reviewer scanning a diff. They have absorbed millions of examples of what "good code" looks like syntactically. What they have not absorbed is what "correct behavior" looks like under concurrent load, with stale caches, during a partial network partition, or when the third-party API returns a 429 instead of a 200.

As Addy Osmani wrote in what may be the most important essay on code review this year: "Code generation became cheap while understanding stayed expensive." Agentic code review means reviewing code whose author cannot explain itself. Classic review validates a colleague's reasoning. Agentic review must reconstruct reasoning that was never written down.

View original post on addyosmani.com → ℹ️ The New Relic number that should scare you: 62% of engineering teams now ship AI-generated code to production without line-by-line manual verification. They trust the confidence of the code itself. The confidence paradox means the code that looks most trustworthy is often the code that fails most surprisingly.

>= to >." The queue grows faster than the team can read.

The math is brutal. Faros AI's telemetry data shows that AI adoption correlates with 98% more PRs that are 154% larger, while review times have grown 91%. Zero-review merges are up 31%. The review queue is now a conveyor belt moving faster than anyone can watch.

This is not a people problem. It is a structural problem. As Tian Pan wrote: "The dominant failure mode of code review in 2026 is that reviewer instincts that worked on human-authored PRs break down on agent PRs because the bugs cluster in different places and the artifacts the reviewer sees are no longer the artifacts that matter."

View discussion on Hacker News → When a human writes a PR, reviewers develop a sense for where bugs hide. They check the boundary conditions, the error paths, the off-by-one opportunities. When an agent writes a PR, the code is syntactically flawless. The bugs are in the assumptions, not the implementation. Reviewers have to develop entirely new instincts, and most have not had time to.

Theo's response to ThePrimeagen's skeptical takes on AI coding agents captures the tension well. ThePrimeagen's core complaints, which Theo largely concedes, include vibe-coded slop PRs as a real and growing review burden, juniors who let the agent do everything never building intuition, and auto-merge tooling that ships without a human gate. Theo pushes back on one point: experienced engineers using agents as power tools genuinely hit a higher productivity ceiling. The question is whether the review infrastructure can keep up.

The comparative study of agentic PRs found another pattern that amplifies review fatigue: agents ghost when they receive subjective feedback. A human contributor adjusts their approach when a reviewer says "this doesn't fit the project's pattern." An agent either ignores the feedback or generates an entirely new PR from scratch. 61.38% of agent-authored PRs carry no recorded review activity at all.

💡 The review fatigue test: Count the unreviewed agent PRs in your org's last sprint. If it is over 20%, you have a structural problem, not a discipline problem. No amount of "review your PRs" Slack reminders will fix a conveyor belt moving faster than humans can watch.

One PR is fine. A hundred PRs is a slow-motion rewrite.

The subtlest horseman does not show up in any single PR. It shows up over weeks and months as agent-authored changes silently shift the architecture. Each change is reasonable in isolation. Together, they constitute a drift that no one approved and no one noticed until the system became unmaintainable.

SoftwareSeni's analysis documents how agents produce contract violations, cross-service dependency breaks, and architectural regressions that pass tests. The tests pass because each change is locally correct. The architecture degrades because no test asserts "the system as a whole still makes sense."

This is the horseman that connects to the 48K files deletion incident we covered recently. Blast radius is the real infrastructure problem. When an agent can make hundreds of changes per day, the cumulative architectural impact dwarfs anything a human team would produce, because each individual change is too small to trigger alarm bells.

View original post on starkravingfinkle.org → The Anthropic trends report notes that developers can "fully delegate" only 0-20% of tasks to agents, yet 41% of code is AI-generated. The gap is filled by supervision that is often cursory. As Andrej Karpathy warned, the coming "slopacolypse" is not code that fails immediately but code that is "almost right, but not quite," degrading system quality gradually.

⚠️ Against the grain: The instinct to hire more reviewers or slow down the merge rate is backwards. Greptile's data shows agent-generated code already produces fewer rewrites than human code (Codex at 5-6% rework vs human baseline of 10%). The problem is not quality — it is the mismatch between what we review for and what actually breaks.

Here is the uncomfortable truth the data supports: agent code fails differently, not more frequently. The MSR study found that only 35.7% of rejected agent PRs reflected genuine agentic failures. 31.2% were rejected for workflow constraints (duplicate PRs, wrong branch, format issues) and 33.1% lacked observable decision rationale. More than half of "agent failures" are actually process failures.

As Simon Willison articulated, the key skill is not review but verification: "being able to confidently instruct agents on how to make changes and then confidently verify that those changes have been applied in the correct way." We explored this distinction in Verify, Don't Review.

The shift is from reviewing code (reading diffs) to verifying behavior (running assertions). Review asks "does this code look correct?" Verification asks "does this code do the correct thing?" The first question depends on human judgment that scales linearly. The second can be automated.

For teams shipping agent-authored PRs today, here is the minimum viable defense: Gate, don't review. Move your quality enforcement from human review to automated gates. Contract tests, architecture fitness functions, property-based tests, and mutation testing catch the four horsemen's failure modes at CI time.

Separate the generators from the judges. The AI that writes the code must not be the AI that reviews it. CodeRabbit, Greptile, and Augment exist because this separation is fundamental, not optional.

Tier your review investment. Agent-generated PRs that only touch files within one module and pass all automated gates need a 2-minute glance. PRs that cross service boundaries or change API contracts need a full human review. Allocate accordingly.

Track architectural drift explicitly. Run coupling metrics, dependency analysis, and module boundary checks weekly, not just per-PR. The horseman you do not see is the one that kills you.

Invest in blast-radius limits. Sandbox the agent's scope per task. An agent working on a login flow should not be able to touch the billing module. We covered the infrastructure for this in It Fails on the Harness, Not the Model.

The four horsemen are not going away. Agent PRs will only increase in volume. But they are survivable, and the teams that build the right verification infrastructure now will have a structural advantage over those who keep trying to read every diff by hand.

Originally published at AgentConn

── more in #ai-agents 4 stories · sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/four-horsemen-of-age…] indexed:0 read:9min 2026-10-03 · —