{"slug": "the-end-is-nigh-for-code-review", "title": "The End Is Nigh for Code Review", "summary": "A July 2026 study from Carnegie Mellon and Stanford researchers, \"AI Writes Faster Than Humans Can Review,\" found that pull request volume grew 3.1× while the reviewer pool grew only 1.5× across 802 developers and nearly 200,000 pull requests at one mid-sized AI-forward company, roughly doubling each reviewer's load. The share of pull requests receiving any written human feedback fell from 39% to 21%, and AI-authored pull requests took about 20% longer to move from first human review to merge. Anthropic reported its typical engineer merged 8× as much code per day in Q2 2026 as in 2024, OpenAI reported a 500% increase in landed pull requests on some teams in the first three weeks of using its Symphony orchestrator, and GitHub went from 1.4 billion to 2.9 billion monthly commits between April and August 2026.", "body_md": "Reviewing every change by hand is ending. For many teams, it already has.\n\nCode is mostly written by agents. This shifts the bottleneck downstream to verification. Code changes can be classified into low-risk vs. elevated-risk changes. Low-risk changes can be auto-reviewed by agentic systems at scale. \n\nIn the physical world, we do this at airports (TSA PreCheck), and it works.\n\n## Code production is accelerating\n\nIn Hemingway’s *The Sun Also Rises*, a character describes going bankrupt as happening \"gradually and then suddenly.\" Code volume is in the suddenly phase, and the growth itself is speeding up. \n\nStart with what the companies building these tools report about themselves. Anthropic says [__the typical engineer there merged 8× as much code per day in Q2 2026 as in 2024__](https://www.anthropic.com/institute/recursive-self-improvement). OpenAI reports [__a 500% increase in landed pull requests (PRs) on some teams__](https://openai.com/index/open-source-codex-orchestration-symphony/) in the first three weeks of using its Symphony orchestrator. Inside OpenAI's research org, [__agents now put in about three workdays of effort for every human workday__](https://openai.com/index/research-acceleration-view-inside-openai/). GitHub went [__from 1.4 billion to 2.9 billion monthly commits__](https://github.blog/news-insights/company-news/the-august-17-outage-and-the-work-ahead/) between April and August 2026. \n\nThose are vendors describing their own products, so discount accordingly. But people outside the vendors are seeing it too. Gergely Orosz, author of *The Pragmatic Engineer*, reported that GitHub saw [__more agent-authored PRs than human-authored ones in August 2026__](https://newsletter.pragmaticengineer.com/p/the-state-of-the-tech-industry-in).\n\n \nAgent-authored PRs have surpassed human-authored ones.\n\n \n## The bottleneck shifted to verification\n\nAnthropic says it has already hit this bottleneck internally: “[__as we've begun to push more code around the organization, human code review has become a new bottleneck.__](https://www.anthropic.com/institute/recursive-self-improvement)”\n\nThe best data I've seen on this is a July 2026 study from researchers at Carnegie Mellon and Stanford, “[__AI Writes Faster Than Humans Can Review__](https://arxiv.org/abs/2607.01904).” They followed 802 developers and nearly 200,000 pull requests at one mid-sized, AI-forward company, with data through April 2026.\n\nPR volume grew 3.1×. The reviewer pool grew 1.5×.\n\n \nPull request volume against reviewer pool growth at one company, through April 2026.\n\n \nDo the math and each reviewer is carrying roughly twice the load they were. AI-authored PRs also took about 20% longer than human PRs to get from first human review to merge, even after controlling for PR size and author.\n\n## Models are getting better, but they still need supervision\n\nSecurity is where the gap shows. In [__a July 2026 benchmark of eight models__](https://arxiv.org/abs/2607.23088), every one had an average vulnerability rate above 56% across real-world risk scenarios like ambiguous requirements and conflicts between security and functionality. That benchmark used older model versions, and newer ones will do better.\n\nBetter models come with their own Jevons paradox for security. They make fewer mistakes but they produce more code. Here's some napkin math (illustrative, not measured):\n\n100 changes × 2% escape rate = 2 defects.\n\n400 changes × 1% escape rate = 4 defects.\n\nYou cut the defect rate in half and doubled the number of defects. Quality improvements get eaten by volume, and volume is growing faster.\n\n## We’re not supervising them well \n\nHere’s the most interesting part of the Carnegie Mellon and Stanford study: what reviewers did with the extra load. \n\nThe share of PRs that got any written human feedback fell from 39% to 21%. Silent approvals per reviewer roughly doubled. Commented reviews stayed flat, at about three a month for the median reviewer. \n\nSo people didn’t review harder. They approved more, and said less. \n\nAn engineer at a mid-sized startup put it bluntly to Orosz: [__“Everyone is playing the theater of doing reviews.”__](https://newsletter.pragmaticengineer.com/p/the-state-of-the-tech-industry-in) \n\n## Beating human review isn’t hard\n\nThe standard objection to AI code review is that it will miss things. True, but so do human code reviews.\n\nAI review doesn’t need to be perfect. It needs to be better than human review. In [__a controlled study presented at the International Conference on Software Engineering in 2022__](https://arxiv.org/abs/2202.04586), human reviewers caught seeded security flaws anywhere from 3% to 61% of the time. Developers doing ordinary code review caught one flaw 3% of the time and the other 21% of the time. Even reviewers told to focus on security, or handed a tailored security checklist, topped out at 61%.\n\n \nDevelopers doing ordinary code review caught two seeded security flaws 3% and 21% of the time. Even with a tailored security checklist, the best result was 61%.\n\n \nThat’s two specific vulnerabilities in one study, not a universal detection rate. But think about how your team reviews code today. You’re closer to 3% than 61%.\n\nAnd humans are better at other things: steering, judgment calls, and the exceptions nobody wrote a rule for. *Save human review for the few changes that matter.*\n\n## One extreme: no mandatory review at all\n\nAmpcode.com, the 20-person team behind the Amp coding agent, doesn't use pull requests. Nobody has to approve a change before it ships. Engineers push straight to main, and Amp ships continuously.\n\nThe usual reaction is that this can't pass a SOC 2 audit. Amp's Will Dollman explains why it does in [__\"That's not SOC 2 compliant\"__](https://ampcode.com/notes/thats-not-soc-2-compliant): \"SOC 2 doesn't require pull requests. It requires that you think about your risks.\"\n\nAmp’s auditor worked with them on controls that fit how they build. Only the right people can push to main. Every commit is signed, so the author is verifiable. Every change runs through automated tests and security checks, and bad changes are blocked. And every commit links back to the conversation that produced it, so the audit trail is as good as a PR's. As Dollman puts it, \"The criteria don't say a second human has to stare at a diff.\"\n\nAmp is also clear about the limits: \"we're not going to pretend a 2,000-person company should let everyone push to main.\" I agree with them.\n\nWhat their setup shows is that human review is one control among several. It can be replaced, if what replaces it is good enough.\n\nAs a larger organization, Cursor also mostly automated their code reviews (according to “people familiar with the matter”), though the bots tag humans occasionally. But they have some of the best AI engineers and unlimited tokens. Most of us are nowhere near Cursor.\n\n## You can’t YOLO at enterprise scale\n\nAt a large company, a bad change can halt operations, break payments, or expose customer data.\n\n[__One in five respondents to the Uptime Institute's latest survey said their most recent major outage cost more than $1 million__](https://uptimeinstitute.com/about-ui/press-releases/uptime-announces-annual-outage-analysis-report-2026). IBM and Ponemon put [__the global average cost of a data breach at $4.99 million__](https://www.ibm.com/reports/data-breach), a record high. \n\nSo you have to decide what the right risk-reward threshold is for your organization. Just remember that playing supersafe is a big long-term risk too. You will be left behind. \n\n## How to get started: let humans handle the exceptions\n\nHere's the operating model I'm proposing. Every change gets reviewed for correctness, security, and alignment with the intended architecture. Routine changes proceed automatically when the context and the verification evidence satisfy policy. Consequential changes go to a human: new service boundaries, major new dependencies, changes to the threat model, and real trade-offs.\n\nIdeally, these human reviews should come with targeted questions to consider. Humans shouldn’t be on the hook for a mega PR just because a review agent couldn’t get all the way to an approval. The human task should be focused on the missing part, not the whole PR.\n\n \nEvery change is reviewed automatically. Routine changes proceed on their own, and consequential changes go to a human.\n\n \nAnthropic’s [__AI-native SDLC playbook__](https://claude.com/blog/the-ai-native-sdlc-playbook) lands in the same place: “layers of agentic review with human review reserved for regulated and critical code.” \n\nYou can get there iteratively. You teach the system where it can act alone and where to pass to a human, and expand that as the evidence supports it:\n\n1. Capture intent. Record service boundaries, dependency rules, critical assets, and what should trigger an escalation.\n2. Compare decisions. Run alongside your reviewers. Investigate every risk it missed and every escalation that wasn't needed.\n3. Enable by class. Approve bounded classes of change. Turn expert decisions into architecture records, tests, and rules so the next one doesn't need the expert.\n\nOpenAI describes the same path in its [__Defense Factory__](https://openai.com/the-defense-factory/): \"We started with small batches and human review, then removed repeated manual steps as the results earned trust.\"\n\nMeasure quality and human effort alongside throughput the whole way. Throughput alone is how you end up with feedback falling from 39% to 21%.\n\n## How we’re doing this at Semgrep\n\nWe’re running this on our own pull requests. It’s a pilot, and I’ll write a follow-up blog to report our learnings at the beginning of next year.\n\nIt starts with routine PRs. Developers opt in (it’s not mandatory, though I think every developer will line up to participate). Agents review the change, and they either approve the PR or hand it off to a human. \n\nBefore the agent approves anything on its own, we define what a safe PR looks like, and our experts label real PRs to use as test cases. Every change runs through four checks (knowledge transfer, product, design, and correctness), and all four have to pass. Then we run the agent in shadow mode beside our human reviewers and track false approvals, coverage, cost, and speed. Where the results hold, we turn it on repo by repo, and each repo sets its own risk level. \n\n## At steady state, fewer than 1% of PRs should need a human\n\nThat’s the goal. Routine changes are approved with their context refreshed, behavior and architecture checked, and evidence recorded. When a change fails its checks, the agent fixes the failure and runs the checks again. After repeated failures, a human takes over. Consequential changes go to humans, with the options and trade-offs already laid out, so people steer the architecture and make the final call.\n\nKeep tracking escaped defects, architectural drift, and human effort.\n\n## The end of reviewing every change by hand\n\nReview every change automatically. Approve what passes. Retry what fails. Humans decide the exceptions, and autonomy grows with evidence. \n\nI'm presenting this at Gartner IT Symposium/Xpo on October 22, 10:00 to 10:20 am, on Stage 3 (Pacific). If you're there, [__come argue with me__](https://semgrep.dev/events/symposium-26-expo-session/). If you can't make it, I'd like to hear how your team is handling review volume. [__Find me on LinkedIn__](https://www.linkedin.com/in/daghanaltas/).", "url": "https://wpnews.pro/news/the-end-is-nigh-for-code-review", "canonical_source": "https://semgrep.dev/blog/2026/end-is-nigh-for-code-review", "published_at": "2026-10-09 00:00:00+00:00", "updated_at": "2026-10-09 15:55:12.096794+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-agents", "ai-tools", "developer-tools", "ai-safety"], "entities": ["Anthropic", "OpenAI", "GitHub", "Carnegie Mellon University", "Stanford University", "Gergely Orosz", "The Pragmatic Engineer", "Symphony"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-end-is-nigh-for-code-review", "markdown": "https://wpnews.pro/news/the-end-is-nigh-for-code-review.md", "text": "https://wpnews.pro/news/the-end-is-nigh-for-code-review.txt", "jsonld": "https://wpnews.pro/news/the-end-is-nigh-for-code-review.jsonld"}}