{"slug": "how-much-does-ai-improve-software-development-productivity", "title": "How Much Does AI Improve Software Development Productivity?", "summary": "A systematic review of 181 documents, including 116 empirical studies, finds that AI-assisted software development accelerates code production more reliably than delivery, with controlled assistant studies reporting roughly 20–30% coding-stage gains and the strongest delivery study observing 10–30% higher release output. The review, covering evidence from 1 January 2022 to 31 July 2026, cautions that faster generation yields business value only when verification, review, and deployment scale accordingly, and recommends measuring productivity as reliable software delivered to production rather than code volume.", "body_md": "## Key findings\n\n- Controlled assistant studies generally find approximately 20–30% coding-stage gains, although results include negative effects (\n[Google RCT](https://arxiv.org/abs/2410.12944);[three field experiments](https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2025.00535)). - The strongest delivery study observes approximately\n[10–30% higher release output](https://www.nber.org/papers/w35275), but non-random adoption creates serious risk of bias. - Agent-native cases report a\n[4.5× median](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/)and an[18× task-specific upper end](https://www.salesforce.com/news/stories/how-engineering-became-agentic/); these low-weight estimates depend on redesigned workflows. - Speed without scaled verification can displace work into review, integration, security and maintenance (\n[longitudinal repository evidence](https://arxiv.org/abs/2511.04427);[DORA 2025](https://dora.dev/research/2025/dora-report/)).\n\n## Abstract\n\n**AI-assisted software development accelerates code production more reliably than it accelerates delivery.** This systematic review of AI coding productivity examines 181 documents published from 1 January 2022, including 116 empirical studies. Forty-four productivity and delivery studies contribute 46 effect estimates; another 72 examine review, quality, security, maintainability, skills and interaction cost. Software development is also a useful—though imperfect—proxy for other AI-assisted business processes because its instructions, revisions and outcomes are digitally recorded.\n\nThe evidence is grouped into three operating models:\n\nthe model suggests; the developer implements and verifies.[Assisted development:](#task-results)the agent implements a defined outcome; a human reads and approves the result.[Spec-driven execution:](#specified-results)agents build, test and perform initial review; humans own intent, architecture and release.[Agent-native delivery:](#agent-native-results)\n\nControlled assistant studies generally report approximately 20–30% coding-stage gains, although the range includes a 19% slowdown. The strongest delivery study observed only **10–30% higher release output** across successive tool generations. The review therefore uses **1.2×** as an observational release-planning case. The larger **2×** spec-driven and **4.5×** agent-native scenarios concern narrower implementation workflows and carry less evidence weight. Faster generation produces business value only when verification, review and deployment scale with it. Productivity should therefore be measured as reliable software delivered to production, not as code volume.\n\n## How was the review conducted?\n\nWe searched for evidence published from **1 January 2022 to 31 July 2026** using three complementary routes: agentic web retrieval and citation chasing, 14 structured OpenAlex searches and direct searches of 15 software-engineering journals. Sources were then screened, appraised and synthesised in three distinct stages.\n\n### How did PRISMA screening determine inclusion?\n\nWe used the [PRISMA 2020 structure](https://www.prisma-statement.org/prisma-2020-flow-diagram) to record identification, deduplication, title-and-abstract screening, full-report assessment and exclusion reasons. The eligibility rules determined whether a source entered the review; PRISMA documented those decisions and counts. It did not grade the credibility of an included study. One AI-assisted reviewer performed the screening, without independent duplicate review. The complete counts and flow are shown in the [PRISMA screening results](#prisma-flow).\n\n### How did risk of bias affect the conclusions?\n\nAfter inclusion, every empirical study received a formal risk-of-bias assessment appropriate to its design: [RoB 2](https://methods.cochrane.org/bias/resources/rob-2-revised-cochrane-risk-bias-tool-randomized-trials) for randomised trials, [ROBINS-I V2](https://www.riskofbias.info/welcome/robins-i-v2) for non-randomised interventions and [JBI tools](https://jbi.global/critical-appraisal-tools) for other designs. Risk of bias was not an automatic exclusion rule. Instead, it reduced the weight placed on studies whose results could also be explained by non-random adoption, team differences, historical baselines or concurrent workflow changes. This is why controlled assistant findings support firmer conclusions than the larger agent-native cases. The design-level judgements are reported in the [risk-of-bias results](#risk-of-bias).\n\n### How was the evidence synthesised?\n\nRisk of bias was combined with outcome directness, precision and relevance to assign each study an evidence-weight grade. Because the studies measured incompatible outcomes and workflows, we did not calculate a pooled meta-analytic effect. Findings were grouped by degree of automation and delivery stage, and ratios were calculated only when the source comparison allowed it. A sensitivity analysis then removed results at serious or high risk of bias to test which conclusions remained. This preserved the approximate 20–30% bounded-assistant finding, but not a precise causal release multiplier or 4.5× agent-native effect.\n\n### Detailed methods, screening and appraisal results\n\n### What question and period did the review cover?\n\nThis systematic review asked: *how does measured software-development productivity change as AI use progresses from prompt assistance to specified delegation and agent-native delivery?*\n\nWe reviewed evidence published from **1 January 2022 to 31 July 2026**. The OpenAlex search ran on **29 July 2026**, followed by a direct search of 15 software-engineering journals on **31 July 2026**. Fourteen archived OpenAlex title-and-abstract queries covered AI coding assistance, agents, productivity, delivery, review, quality, security, maintainability, skills and developer interaction. Citation chasing supplemented the database and journal searches with peer-reviewed papers, preprints, official research publications and first-party engineering reports. Where available, a later peer-reviewed report replaced its preprint.\n\n### What evidence was eligible?\n\nProductivity studies had to measure a development workflow against a baseline or comparator. Studies of review, quality, security, maintainability, skills and interaction were included as adjacent evidence. Surveys and qualitative accounts provided context only; unverifiable claims were excluded. Traceability and corrections follow Reinvently’s [research and editorial standards](/research-standards/).\n\nWe kept time and output separate: **25% less time** implies up to **1.33× potential throughput**, while **25% more output** is **1.25×**. Commits and lines of code were treated as activity measures, not production value.\n\n### How was the evidence base divided?\n\nThe **181 source documents** were divided as follows:\n\n**181 source documents****44 productivity and delivery studies**→** 46 effect estimates****72 adjacent-outcome studies****4 secondary syntheses****61 supporting or contextual documents**\n\n#### 1. Which records passed PRISMA eligibility screening?\n\nWe used three retrieval streams. Google searches and citation chasing produced a 115-document agentic-web corpus. Fourteen structured OpenAlex searches then tested and expanded it. Finally, direct Crossref searches enumerated publications from 15 named software-engineering journals. Retrieval route and publication type were recorded separately: a journal article first found through the agentic search remains in that stream, while the later journal match is recorded as a duplicate. We archived search results, deduplication, screening decisions and exclusion reasons using the [PRISMA 2020 updated-review structure](https://www.prisma-statement.org/prisma-2020-flow-diagram). One AI-assisted reviewer screened the records; screening was not independently duplicated.\n\n| Direct journal source | Records | Topical | Assessed | Corpus contribution |\n|---|---|---|---|---|\n| Empirical Software Engineering | 859 | 70 | 6 | 6 new; 3 existing duplicates |\n| IEEE Transactions on Software Engineering | 1,145 | 58 | 4 | 3 new; 2 existing duplicates |\n| ACM Transactions on Software Engineering and Methodology | 1,238 | 311 | 9 | 8 new; 2 existing duplicates; 1 replacement |\n| Journal of Systems and Software | 1,205 | 36 | 3 | 3 new |\n| Information and Software Technology | 980 | 31 | 4 | 4 new |\n| IEEE Software | 978 | 43 | 1 | 1 new |\n| Automated Software Engineering | 331 | 35 | 0 | — |\n| Software Quality Journal | 197 | 9 | 1 | 1 new |\n| Journal of Software: Evolution and Process | 514 | 19 | 0 | — |\n| Software: Practice and Experience | 475 | 12 | 1 | — |\n| Science of Computer Programming | 480 | 3 | 1 | 1 new |\n| Requirements Engineering | 97 | 8 | 0 | — |\n| ACM Transactions on Computing Education | 292 | 16 | 0 | — |\n| IEEE Transactions on Dependable and Secure Computing | 2,000 | 3 | 0 | — |\n| ACM Transactions on Privacy and Security | 204 | 4 | 0 | — |\nTotal | 10,995 | 658 | 30 | 27 new; 10 existing duplicates; 1 replacement |\n\n## Show the search terms\n\n**Agentic Google searches, normalised by topic:** AI coding productivity; AI coding assistant productivity study; GitHub Copilot productivity study; AI-assisted software development productivity; coding agent productivity; agentic coding productivity; spec-driven development; AI coding release output; deployment velocity; software delivery lead time; AI code review productivity; AI-generated code quality; AI-generated code bugs; AI-generated code maintainability; AI-generated code security vulnerabilities; AI coding greenfield brownfield; AI coding developer experience seniority; AI coding programming languages static typing; Stack Overflow AI developer survey; and first-party research from Google, Microsoft, GitHub, Anthropic, Meta and AWS.\n\n**Publication hosts within the agentic-web stream:** arXiv, 40 documents; ACM Digital Library or ACM DOI records, 14 documents; and other web, publisher or repository sources, 61 documents. These are host or publisher categories, not additional retrieval streams; the ACM group includes both journal and conference publications.\n\n**Exact OpenAlex queries:** AI coding assistant; AI code assistant; AI-assisted programming; AI-assisted software development; generative AI software development; coding agent; agentic coding; agentic code review; GitHub Copilot productivity; software development productivity AI; AI code review; AI generated code maintainability; AI generated code security; and spec-driven development.\n\n**Direct journal search:** 15 journals were enumerated by ISSN for articles published from 1 January 2022 to 31 July 2026. Ten reports already present in the included corpus were removed by DOI or normalised title. Five further records overlapped the earlier OpenAlex screen; their richer publisher records were reassessed rather than counted as new identifications. The complete [journal search log](/research/ai-coding-productivity/prisma/journal-search-log.csv), [screening register](/research/ai-coding-productivity/prisma/journal-screening-register.csv) and [report assessments](/research/ai-coding-productivity/prisma/journal-report-assessments.csv) are published with the review.\n\nThe earlier verbatim Google query history was not retained, so the first list records the normalised concepts rather than claiming an exact query log. The [OpenAlex query log](/research/ai-coding-productivity/prisma/search-log.csv) is exact.\n\n#### 2. What did the risk-of-bias assessment find?\n\nRisk of bias tests whether a study’s design could systematically overstate or understate its result. It does not measure the size, relevance or practical usefulness of the finding. A large observational study may therefore provide useful scenario evidence even when it cannot establish that AI caused the reported effect.\n\nThe clearest causal evidence came from randomised trials. Most release-level and agent-native findings came from non-randomised studies, cases or benchmarks, where team differences, adoption choices and concurrent workflow changes could explain part of the result.\n\n| Study design | Studies | Risk judgements | What this means |\n|---|---|---|---|\n| Randomised trials | 10 | 9 some concerns; 1 high | Best causal evidence, with some reporting or adherence limitations |\n| Non-randomised interventions | 18 | 2 moderate; 16 serious | Other differences between adopters and controls may explain part of the effect |\n| Quasi-experiments | 18 | 13 some concerns; 5 high | Comparison groups were not always equivalent |\n| Observational studies | 39 | 28 some concerns; 11 high | Useful for associations, but not isolated causal effects |\n| Cases and benchmarks | 31 | 15 some concerns; 16 high | Useful for showing possibilities and failure modes, not expected gains |\n\nNone met the strict low-risk standard across every applicable domain. This does not make the evidence unusable; it means precise causal claims require caution. Randomised studies were assessed with [RoB 2](https://methods.cochrane.org/bias/resources/rob-2-revised-cochrane-risk-bias-tool-randomized-trials), non-randomised interventions with [ROBINS-I V2](https://www.riskofbias.info/welcome/robins-i-v2), and other designs with [JBI tools](https://jbi.global/critical-appraisal-tools). The [116 study-level assessments](/research/ai-coding-productivity/risk-of-bias/risk-of-bias-register.csv), [summary](/research/ai-coding-productivity/risk-of-bias/risk-of-bias-summary.json) and [decision rules](/research/ai-coding-productivity/risk-of-bias/README.md) are published with the review.\n\n#### 3. How was each study weighted?\n\nEach of the **116 empirical studies** received an evidence-weight grade. This incorporated the risk-of-bias judgement and also considered whether the outcome directly answered the review question, how precisely it was measured and how relevant the workflow was to real software delivery.\n\n| Evidence weight | Included studies | Interpretation |\n|---|---|---|\n| High | 12 | Credible comparator, sufficiently direct outcome and no identified material limitation likely to reverse the result |\n| Moderate | 77 | Useful comparative or directly measured evidence with residual bias, confounding, indirectness or precision limitations |\n| Low | 27 | Small or weakly controlled study, historical baseline, self-report, organisational case or counterfactual estimate |\n\nContextual documents and secondary syntheses were not graded as primary evidence. The studies were too different for a meaningful pooled estimate, so findings were grouped by degree of automation and delivery stage. Causal claims were limited to credible comparisons; telemetry and case estimates were treated as associations. Every productivity conclusion can be traced from its **effect estimate ( E)** to its\n\n**study (** and source. The\n\n`ST`\n\n)[study-level evidence-weight register](/research/ai-coding-productivity/evidence-weight-register.csv)publishes one grade for each of the 116 empirical studies, including all 72 adjacent-outcome studies. The separate\n\n[operating-model audit](/research/ai-coding-productivity/operating-model-audit.csv)retains the effect-level grades and plotting decisions for the 46 productivity estimates.\n\n#### 4. What survived the sensitivity analysis?\n\nRemoving results with serious or high risk of bias still leaves bounded assistant studies pointing to gains of roughly **20–30%**. It leaves insufficient evidence for a precise release multiplier or a causal **4.5×** agent-native gain. The **1.2×** and **4.5×** figures are therefore planning scenarios, not expected effects.\n\n## How do reported gains vary by degree of automation?\n\nThe findings are grouped by the three operating models defined in the abstract, each representing a different degree of automation. A separate downstream-conversion section then tests how gains in coding activity survive into releases and actual use.\n\n### 1. Assisted development: task completion and local output\n\nThe strongest evidence concerns assistants operating inside otherwise conventional human-led workflows. The measured effects are heterogeneous, and tool choice changes the interaction model without determining the outcome by itself. For a practical comparison of those interaction models, see [GitHub Copilot vs Claude Code vs Cursor](/blog/github-copilot-vs-claude-code-vs-cursor/). Teams translating the evidence into a budget can use the [AI coding tool cost calculator](/tools/ai-cost-calculator/) to model assisted, spec-driven and agent-native usage separately.\n\nA randomised trial involving 96 Google engineers found that AI assistance reduced time on a complex enterprise task by an estimated [21%](https://arxiv.org/abs/2410.12944), although the confidence interval was wide. Three field experiments at Microsoft, Accenture and a Fortune 100 company covered 4,867 developers and found a [26.08% increase in completed tasks](https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2025.00535).\n\nA 109-participant controlled experiment found that ChatGPT helped with simple coding puzzles but [did not improve efficiency or quality on a more typical software-development task](https://arxiv.org/abs/2402.05650). A later controlled study of 69 participants found no strong overall time effect, but AI users achieved [more than twice the median task completeness](https://doi.org/10.1109/TSE.2026.3679627) and scored 12.5% lower on code-ownership questions. Together these results show why task completion, time and understanding should be reported separately.\n\n**Bottom line:** multiple controlled studies support an assisted-development reference value near **1.25×**, but the effect varies by task, repository familiarity and developer experience. The range includes null results and a 19% slowdown. This is the strongest of the three evidence tiers.\n\n## Read the extended assisted-development evidence\n\nTwo smaller crossover experiments add qualified positive evidence. In a 32-person ICSE study, an in-IDE code-understanding assistant helped participants complete [0.47 more subtasks than web search](https://research.google/pubs/using-an-llm-to-help-with-code-understanding/), but did not significantly change completion time or comprehension. An 18-person JavaScript study found that participants using ClueBot [finished faster and made fewer errors](https://doi.org/10.1145/3702163.3702168); its small student-heavy sample and high risk of bias limit generalisation.\n\nThe updated search adds two comparative task studies without changing that conclusion. In a 27-student controlled study, three completion tools reduced average recorded time by [11–14%](https://arxiv.org/abs/2404.12000), but only two of the seven participants who finished both assigned tasks were faster with AI. A quasi-experiment found that some non-programmers using ChatGPT completed small, well-specified tasks faster than programmers without AI, but the design lacked both a programmers-with-AI group and a non-programmers-without-AI group ([Esnaola et al.](https://doi.org/10.1007/s10664-026-10813-7)).\n\nA larger preregistered maintainability experiment also captured an assisted-development speed result. In its first phase, participants who chose an AI assistant completed a Java feature task in [30.7% less median time](https://link.springer.com/article/10.1007/s10664-026-10889-1). The comparison was observational rather than randomised, but the study explicitly predates autonomous coding agents and therefore belongs at Level 1.\n\nA larger quasi-experiment at Ant Group compared matched teams during the staged introduction of CodeFuse. Across a sample of 1,219 programmers, code output increased by [55%](https://www.bis.org/publ/work1208.pdf). More conservative specifications found gains of roughly 13–22% in tasks completed and related workflow measures. Effects were concentrated among junior developers, and lines of code remained the principal outcome.\n\nA public-sector study by GovTech Singapore reported [21–28% faster coding and task completion](https://arxiv.org/abs/2409.17434) with GitHub Copilot. The direction is consistent with the larger controlled studies, although the disclosed design provides less protection against selection and reporting effects.\n\nA separate Anthropic randomised trial found only a small, statistically non-significant task-speed improvement, while participants using AI scored [17 percentage points lower on a subsequent mastery test](https://www.anthropic.com/research/AI-assistance-coding-skills) (50% against 67%). This does not establish a long-term productivity loss, but it identifies skill retention as an omitted outcome in most speed studies.\n\nEnterprise studies based mainly on self-report provide useful triangulation but weaker effect estimates. An [IBM study of 669 developers](https://research.ibm.com/publications/examining-the-use-and-impact-of-an-ai-code-assistant-on-developer-productivity-and-experience-in-the-enterprise) found heterogeneous perceived productivity effects from watsonx Code Assistant, while a Microsoft randomised diary study found that [84% reported positive changes in daily work practices](https://www.microsoft.com/en-us/research/publication/dear-diary-a-randomized-controlled-trial-of-generative-ai-coding-tools-in-the-workplace/) without a corresponding change in trust in generated code.\n\nInteraction studies expose work that completion metrics miss: prompting, inspecting, editing and recovering from suggestions. Microsoft’s [CUPS study of 21 programmers](https://www.microsoft.com/en-us/research/?p=894132) made those costs observable. A randomised study of [proactive programming assistance](https://www.microsoft.com/en-us/research/publication/need-help-designing-proactive-ai-assistants-for-programming/) found benefits but strong interface effects. A five-day professional field study similarly found that poorly timed proactive assistance could interrupt flow ([ProAIDE](https://doi.org/10.1145/3742413.3789148)). Across 66,239 industrial interactions, accepted suggestions differed systematically by developer history, project history and code context; a predictive filter reduced likely interruptions, but did not estimate delivery speed ([suggestion-acceptance study](https://doi.org/10.1145/3808125)).\n\nNew journal evidence reinforces that interaction design changes outcomes without supplying an end-to-end multiplier. In a data-science user study, naming the current workflow step significantly increased recommendation acceptance, while presenting alternative code paths did not ([ST-104](https://doi.org/10.1007/s10664-025-10622-4)). A 27-student instrumented study found acceptance varied by task and suggestion type, and that editing existing code often caused backtracking ([ST-102](https://doi.org/10.1145/3785479)).\n\nAdoption evidence helps explain these mechanisms but does not establish delivery gains. GitHub’s telemetry-linked study of 2,047 respondents associated higher completion acceptance with higher perceived productivity ([Measuring GitHub Copilot’s Impact on Productivity](https://doi.org/10.1145/3633453)). A peer-reviewed [410-developer usability survey](https://doi.org/10.1145/3597503.3608128) identified fewer keystrokes, faster completion and syntax recall as leading benefits, but control and non-functional requirements as common failures. A 481-programmer journal survey found that 84.2% used assistants at least occasionally, while inaccurate output, lack of trust and weak project context remained leading barriers ([Sergeyuk et al.](https://doi.org/10.1016/j.infsof.2024.107610)). A seven-company European case study found use was still predominantly personal and assistive rather than process-level ([Kemell et al.](https://doi.org/10.1016/j.infsof.2025.107805)).\n\nSmaller controlled studies reinforce the importance of task selection. At ANZ Bank, a six-week controlled study reported [42.3% less completion time](https://arxiv.org/abs/2402.05636) on Python exercises. A 24-person within-subject study found higher requirements completed per minute with both conversational and completion interfaces, although its tasks were short and only nine participants were professionals ([Weber et al.](https://doi.org/10.1145/3661145)).\n\nIndustrial completion systems provide credible day-to-day activity measures, although these are not product outcomes. Meta’s CodeCompose supplied 8% of accepted code for 16,000 users ([FSE 2024 deployment study](https://doi.org/10.1145/3643774)); extending it to multi-line suggestions nearly doubled keystrokes saved from 9% to 17% ([large-scale experiment](https://arxiv.org/abs/2402.04141)). The UK Government’s trial made 2,500 licences available and combined telemetry with surveys; participants reported approximately [one hour saved per working day](https://www.gov.uk/government/publications/ai-coding-assistant-trial/ai-coding-assistant-trial-uk-public-sector-findings-report) (56 minutes), but the result was self-reported rather than a causal time study.\n\n### 2. Spec-driven execution: bounded delegation\n\nAt the second level, the engineer defines an outcome, constraints and acceptance checks; the agent performs multi-file implementation; the engineer still reads and owns the change. Evidence here is less controlled because organisations usually adopt several changes at once. [How to Choose an AI Development Framework](/blog/ai-dev-workflow-frameworks-gsd-bmad-openspec-speckit/) compares the main frameworks used to structure this kind of specification and delegation.\n\nTwo bounded comparisons fit this operating model more closely than assisted development. In a 24-developer brownfield-onboarding experiment, Copilot Agent reduced mean completion time by [61.7% relative to Copilot Ask](https://aisel.aisnet.org/pacis2026/ai_fow/ai_fow/13/) and workload by 57.4%, without a significant correctness improvement; interaction shifted from active collaboration towards passive supervision. In a repeated single-developer comparison, Cursor completed the same defined web app in about 13 minutes with two to three prompts, against 27 minutes and five to six prompts for ChatGPT, but there was no manual baseline ([Afungchwi and Rhioui](http://www.theseus.fi/handle/10024/923044)).\n\nA separate controlled comparison found that agents reduced effort and completed some otherwise inaccessible tasks, but did not remove the developer’s need to understand and verify their behaviour ([Code with Me or for Me?](https://arxiv.org/abs/2507.08149)). This supports the workflow distinction without supplying a comparable delivery multiplier.\n\n## Read the extended spec-driven evidence\n\nA peer-reviewed test-driven workflow supplies direct evidence for specification as a verification aid. In a 15-programmer study, TiCoder used generated tests to clarify intent; participants were significantly more likely to evaluate generated code correctly and reported lower task-induced cognitive load ([ST-111](https://doi.org/10.1109/TSE.2024.3428972)). The accompanying model benchmark improved code-generation accuracy, but the human study did not estimate delivery throughput.\n\nAn Amazon Prime Video Financial Systems team provides a more specific example. A senior engineer spent three weeks decomposing work and writing detailed requirements before a ten-day sprint. The team used spec-driven development for complex features and direct agent assistance where requirements were already clear. It reduced a 90-week estimate to 24 weeks—approximately [3.75× acceleration](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/), which AWS rounds to 4×—and produced 556 commits against a baseline of 96, nearly six times the usual output.\n\nThe study conditions limit external validity: six engineers had no on-call work, no other projects, minimal meetings and no context switching. Specification was part of the system, not the only intervention.\n\nCisco reports teams achieving up to [3× productivity](https://www.sonarsource.com/blog/ai-first-engineering-cisco/) in technical-debt workflows. Its agent investigates a quality issue, produces a written plan, implements in a fresh session and opens a pull request for human review. The work is unusually deterministic, with SonarQube providing a ready-made issue description and verification layer.\n\nAmazon reports a broader but similarly bounded transformation programme. Developers review and refine Amazon Q Developer’s plan before its transformation agent implements multi-step upgrades. Across tens of thousands of production Java applications, Amazon estimates that this saved [4,500 developer-years](https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone/). The estimate derives from migrated dependencies and assumed manual effort rather than a concurrent control group, so it demonstrates scale rather than a transferable acceleration factor.\n\nEvidence on specification quality shows why “more detail” is not sufficient by itself. Copilot generated both repayments and reproductions of technical debt when prompted with 1,140 variants derived from TODO comments ([O’Brien et al.](https://lab-design.github.io/papers/ICSE-24b/)). By contrast, Meta found that pre-computing context across 4,100 files reduced agent tool calls by [40%](https://engineering.fb.com/2026/04/06/developer-tools/how-meta-used-ai-to-map-tribal-knowledge-in-large-scale-data-pipelines/), while a separate internal agent reduced approximately ten hours of performance-regression investigation to 30 minutes ([capacity-efficiency case](https://engineering.fb.com/2026/04/16/developer-tools/capacity-efficiency-at-meta-how-unified-ai-agents-optimize-performance-at-hyperscale/)). These first-party cases suggest that explicit context and executable feedback reduce search cost; they do not isolate specification as a causal multiplier.\n\nSpecification also cannot recover business knowledge that was never made explicit. In a controlled .NET study, an agent could produce technically valid changes while violating intended behaviour when historical terminology and business rules were ambiguous ([domain-ambiguity study](http://urn.kb.se/resolve?urn=urn:nbn:se:kau:diva-110473)). An end-to-end meeting-to-pull-request pipeline performed reliably on individual issues, but 44 workflow runs exposed different conflict modes when agents operated in parallel on sparse and structured repositories ([pipeline evaluation](http://urn.kb.se/resolve?urn=urn:nbn:se:miun:diva-57923)). Strong specifications therefore need explicit domain definitions, repository contracts and conflict handling—not task prose alone.\n\nA peer-reviewed ICSE 2026 study examined an expert-led, AI-implemented workflow on PicoScenes, a 1.52-million-line industrial system with a ten-year history. The authors report [68.3% less feature-implementation time](https://doi.org/10.1145/3786583.3786872), alongside 28.2% lower mean cyclomatic complexity and 50% fewer defects. The comparison concerns one system and a bundled development method, but it provides unusually detailed evidence that specified delegation can extend beyond small greenfield tasks.\n\nAnthropic’s internal mixed-methods study offers a second organisational data point. Engineers self-reported a 50% average productivity gain as Claude use expanded, while internal telemetry showed a [67% increase in merged pull requests per engineer per day](https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic). The before-and-after design cannot isolate Claude from team growth, task selection or workflow change.\n\n**Review finding:** reported outcomes support a **1.5–2.5× planning range** for bounded spec-driven execution with human review. The **2×** midpoint is not a pooled effect: one small controlled agent-versus-assistant comparison implies 2.61× potential throughput, while the remaining evidence is heterogeneous and mainly case-based. **Basis:** `E-05`\n\n, `E-21`\n\n, `E-30`\n\n, `E-38`\n\n, `E-39`\n\n, `E-42`\n\n, `E-43`\n\n.\n\n### 3. Agent-native delivery: production cases\n\nThe largest claims concern teams that restructure repositories, planning, testing, review and parallel execution so agents can operate for longer with less synchronous human attention. Where that design extends to several specialised agents, [Multi-Agent Orchestration Frameworks Compared](/blog/multi-agent-orchestration-frameworks-compared/) maps the main architectural categories and their reliability trade-offs.\n\nSalesforce reports that autonomous tools now write code, review pull requests, generate tests and drive deployments across its engineering organisation. Work items completed per developer rose [50.8% year over year](https://www.salesforce.com/news/stories/how-engineering-became-agentic/), while its “Effective Output” measure rose 151.3%. The organisation also standardised on Claude Code, removed token limits and rebuilt workflows, so the result is an agent-native organisational bundle rather than an isolated tool effect.\n\nAWS studied more than 50 teams and found that the 25 combining new tools with new practices outperformed those that simply added AI to existing workflows. In separate Amazon Stores pilots—typical teams working their regular backlogs, with no special conditions or handpicked engineers—the median [productivity gain was 4.5×](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/), with some teams exceeding 10× normalised deployment velocity. Its Bedrock pathfinder reported 20× normalised commit velocity, but this involved six senior engineers rebuilding an inference engine against a 30-developer, 12-to-18-month estimate under highly unusual conditions. Deployment velocity is the more useful number.\n\nOpenAI describes an internal product built with no manually written code: approximately one million lines and 1,500 merged pull requests directed initially by three engineers over five months. The team estimates it took [one-tenth of the manual development time](https://openai.com/index/harness-engineering/). This is a serious production account with active users, but the 10× figure remains the team’s counterfactual estimate rather than an experiment.\n\nSalesforce reports a 33-endpoint migration completed in 13 days rather than an estimated 231 person-days—an [18× task-specific acceleration](https://www.salesforce.com/news/stories/how-engineering-became-agentic/). It used reference implementations, reusable rules, isolated environments and autonomous build-fix-validate loops. Migration work is repetitive and parallelisable, limiting the result’s generalisability to novel product development.\n\n## Read the extended agent-native evidence\n\nUsage telemetry also shows that agentic work remains strongly shaped by human expertise. Anthropic’s analysis of approximately [400,000 Claude Code sessions](https://www.anthropic.com/research/claude-code-expertise) found that people made most planning decisions while the agent made most execution decisions; greater domain expertise was associated with a higher probability of verifiable success.\n\nRepository evidence shows broad adoption without establishing productivity. A study of 128,018 GitHub projects estimated coding-agent adoption at [22.20–28.66%](https://doi.org/10.1145/3822180) by February 2026 and found that agent-assisted commits were larger than human-only commits. Across 26,760 agent-authored pull requests, agents imported libraries in 29.5% of changes but introduced new dependencies in only 1.3%; 75% of those additions specified a version ([library-usage study](https://kclpure.kcl.ac.uk/portal/en/publications/9491a361-b84f-4d93-8b55-a7ed5cca4878)). These traces establish that agents participate in real repository work, but larger changes and dependency interactions also expand the integration surface.\n\nAgent benchmarks define a capability ceiling rather than a workplace-productivity effect. [SWE-Lancer](https://openai.com/index/swe-lancer/) maps more than 1,400 real freelance tasks to approximately $1 million of historical payouts. The [ICLR 2024 SWE-bench study](https://proceedings.iclr.cc/paper_files/paper/2024/hash/edac78c3e300629acfe6cbe9ca88fb84-Abstract-Conference.html) and [SWE-agent](https://arxiv.org/abs/2405.15793) show that repository tools and execution feedback can materially improve issue resolution. [SWE-bench Multimodal](https://arxiv.org/abs/2410.03859) exposes weaker performance on visual JavaScript tasks, while [GitTaskBench](https://arxiv.org/abs/2508.18993) adds workflow-driven repository use and an economic metric. The paired [SWE-Skills-Bench](https://arxiv.org/abs/2603.15401) result is a useful constraint: 39 of 49 packaged skills produced no pass-rate gain and average improvement was only 1.2%. None of these results can be converted directly into hours saved without a human workflow comparator.\n\n**Review finding:** reported agent-native workflows cluster around **2.5–10×**, with a higher task-specific migration estimate. AWS’s **4.5× median** is the most suitable central reference in the available case evidence. Evidence weight remains low because the studies are observational, heterogeneous and frequently vendor-published. **Basis:** `E-29`\n\n, `E-31`\n\n, `E-40`\n\n, `E-41`\n\n.\n\n### Study results by operating model and exposure\n\nThe evidence changes in both magnitude and strength across the three operating models. The graph represents all 34 estimates that can be expressed numerically against a baseline or comparator. Studies were assigned by their observed workflow, not the product name; the [46-estimate classification audit](/research/ai-coding-productivity/operating-model-audit.csv) records each decision. Seven estimates aggregate several modes or do not disclose enough workflow detail, so they appear in a separate cross-level group rather than being forced into an automation tier. `E-11`\n\nappears at all three levels because it reports separate release effects for autocomplete, interactive agents and autonomous agents. Twelve other productivity estimates report directional, absolute or otherwise non-comparable outcomes. Where a study reports time saved, the value is converted to potential throughput using the reciprocal relationship defined under [outcome definitions](#outcome-definitions). Every label resolves to the corresponding [ E entry](#effect-register). The values remain heterogeneous and are not pooled.\n\n## Do coding gains survive into production?\n\nThis conversion is a cross-level outcome rather than a fourth operating model. At Level 1, local task gains must survive integration and review. Levels 2 and 3 can produce much larger increases in implementation activity, but place correspondingly greater demand on validation, release and adoption capacity.\n\nThe largest direct study of that conversion followed more than 100,000 GitHub developers using matched event studies and internal usage telemetry. Autocomplete, interactive agents and autonomous agents were associated with cumulative commit gains of 40%, 140% and 180%, respectively. The final effect was much smaller: the 180% commit increase became [50% more projects and 30% more releases](https://www.nber.org/papers/w35275), while analysis across four application marketplaces found no increase in total usage. This is the clearest evidence that code activity, shipped software and realised value are different outcomes.\n\n## Read the extended delivery and adoption evidence\n\nMETR’s randomised study provides the strongest negative result, but not a clean Level 1 estimate. Sixteen experienced open-source developers could choose among early-2025 chat and agent-capable tools while completing 246 real tasks in familiar repositories. They were [19% slower](https://arxiv.org/abs/2507.09089). Because the result was not separated by interaction mode, it belongs in the cross-level group.\n\nTwo unusually large studies find smaller average effects and strong experience differences. A peer-reviewed analysis in *Science* detected AI-supported Python functions across more than 30 million contributions from 160,097 developers. Within-developer models estimated a [3.6% increase in quarterly output](https://doi.org/10.1126/science.adz9311) at observed adoption levels, concentrated among experienced developers; early-career developers showed no significant gain. The generating workflow was inferred rather than observed, so this result is also cross-level. A regression-discontinuity natural experiment covering 187,489 developers found that Copilot access increased the relative share of coding activity by [12.37%](https://mackinstitute.wharton.upenn.edu/wp-content/uploads/2025/04/WTIC.2025_Nagle-Frank_Generative-AI.pdf) and reduced project-management activity by 24.93%. The latter measures work allocation rather than total delivered output, but provides strong evidence that AI changes the composition of work.\n\nOther quasi-experiments reinforce the same pattern. A generalized synthetic-control study estimated [5.9% higher project-level code contributions](https://arxiv.org/abs/2410.02091) and no code-quality change after Copilot adoption, but coordination time increased by 8%. The preprint’s first version reported 6.5% and 41.6%; the authors revised both figures downwards, and this review follows the current version. A 1,350-developer difference-in-differences study found approximately [12% more commits and 6% more active repositories](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5100609), alongside a 2% rise in copy-pasted code. A workplace study using a time discontinuity estimated [6.9–8.4% higher coding productivity](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6515379) and 3.1–4.3% fewer working hours without a measured quality decline; its “AI coder” workflow is not described precisely enough to assign an automation level. Finally, a Python-versus-R natural experiment estimated [17.82% more package releases and 51% more commits](https://arxiv.org/abs/2409.08379) in the Copilot-supported ecosystem, with gains weighted toward maintenance; differences between language communities remain an important identification caveat.\n\nMatched and longitudinal workplace studies are less uniformly positive. Uplevel compared approximately 800 developers with and without Copilot access and found [no significant change in pull-request cycle time or throughput](https://uplevelteam.com/blog/genai-developers), alongside a 41% higher bug rate. A two-year NAV study covering 26,317 commits across 703 repositories similarly found [no statistically significant post-adoption activity change](https://arxiv.org/abs/2509.20353). By contrast, a 13-month study of three agile teams found one stable team increased delivered story points by [59.1%](https://arxiv.org/abs/2602.13766), with the other teams and quality measures more mixed.\n\nVendor telemetry shows that more code can coexist with more downstream work. Faros’s 2025 dataset of 10,000 developers and 1,255 teams associated high AI adoption with [21% more completed tasks and 98% more merged pull requests](https://www.faros.ai/blog/ai-software-engineering), but also 91% longer review time and no organisation-level DORA improvement. Its later two-year study of 22,000 developers and 4,000 teams found [33.7% more completed tasks and a 16.2% higher pull-request merge rate](https://www.faros.ai/research/ai-acceleration-whiplash), alongside 861% more code churn, 242.7% more incidents per pull request and a fivefold rise in median review time. Both reports aggregate autocomplete, chat and agentic use, so their effects cannot be assigned to one operating model. They also remain observational and are published by an engineering-intelligence vendor.\n\nOther large datasets point in the same two-sided direction. DX reports, from 135,000 developers in 425 organisations, an average [3.6 hours saved per week](https://getdx.com/report/ai-assisted-engineering-q4-impact-report/) and 60% more pull requests among daily users. Jellyfish reports that the highest-adoption cohort across more than 200,000 engineers shipped [56% more epics per 100 engineers](https://jellyfish.co/blog/are-ai-coding-tools-making-companies-more-productive/) and used 25% fewer person-months per epic. These aggregate coding-assistant and agent use rather than reporting mode-specific effects. Neither comparison randomly assigned AI use, so selection, team maturity and task mix remain plausible explanations.\n\nDeveloper surveys explain why perception and telemetry diverge. In Stack Overflow’s [2024 survey, 81% named productivity as a benefit](https://survey.stackoverflow.co/2024/ai). By [2025, 69% of agent users reported increased productivity](https://survey.stackoverflow.co/2025/ai), yet 46% distrusted AI accuracy and 45% said debugging AI-generated code was time-consuming. These are valuable population signals, not estimates of causal speed.\n\nLarge first-party surveys provide adoption context rather than causal estimates. GitHub’s [2024 international survey](https://github.blog/news-insights/research/survey-ai-wave-grows/) documents perceived effects on flow, cognitive effort and collaboration. [JetBrains’ 2025 Developer Ecosystem survey](https://devecosystem-2025.jetbrains.com/) and [Atlassian’s 3,500-person developer-experience survey](https://www.atlassian.com/blog/developer/developer-experience-report-2025) broaden that picture. Anthropic’s analysis of [500,000 coding interactions](https://www.anthropic.com/research/impact-software-development) found that Claude Code use was predominantly automation rather than augmentation. None of these designs establishes an output multiplier.\n\nPeer-reviewed adoption research points to workflow fit rather than tool access alone. A convergent mixed-methods study — 100 engineers surveyed, then 183 more to validate the resulting model — found that [compatibility with existing development workflows](https://doi.org/10.1145/3652154) was the main adoption driver. A survey of 188 engineers instead highlighted [habit and expected performance](https://doi.org/10.1145/3725529), while a 305-participant study found different adoption and verification behaviours between students and professionals ([empirical adoption study](https://doi.org/10.1016/j.infsof.2026.108036)). A separate [South Korean industry survey](https://doi.org/10.1007/s10664-025-10730-1) examined what sustains use after initial adoption. Organisational evidence reaches a similar boundary: a seven-company study found tools still operating mainly as [personal assistants](https://doi.org/10.1016/j.infsof.2025.107805), and a survey of 481 programmers documented broad use alongside persistent technical and organisational barriers ([practice study](https://doi.org/10.1016/j.infsof.2024.107610)). An IBM survey of 57 enterprise developers likewise treated [readiness and workflow requirements](https://doi.org/10.1145/3786181.3788727) as separate from reported benefits. These studies explain adoption conditions; they do not establish causal productivity gains.\n\nThe same distinction matters outside engineering telemetry: an organisation can report faster local work without obtaining adoption or financial value. [The Enterprise AI Reality Check](/blog/enterprise-ai-adoption-reality-check-2026/) examines that wider gap between tool deployment, workflow change, measurement and realised return.\n\n**Review finding:** the controlled evidence centres near **1.25× output**, with a range that includes negative effects. Large comparative studies then show substantial attenuation between code activity and shipped value. In the largest matched event study, release output was approximately **10% higher for autocomplete, 20% higher for interactive agents and 30% higher for the cumulative autonomous-agent stack**. For budgeting and time-to-market analysis, **10–30% higher release output, centred on 20%, is the more defensible base range**; 2–4.5× remains implementation-stage upside requiring workflow redesign and independent verification. Task fit, repository maturity, developer familiarity and downstream capacity explain more of the variation than tool availability alone. **Basis:** `E-01`\n\n–`E-20`\n\n, `E-23`\n\n–`E-28`\n\n, `E-32`\n\n–`E-37`\n\n, `E-43`\n\n–`E-46`\n\n.\n\n## What changes the productivity result?\n\nThe degree of automation does not determine the result by itself. Project context, programming-language feedback and developer experience change both the attainable gain and the amount of verification required.\n\n### Greenfield and brownfield development\n\nProject context is a major source of heterogeneity. Greenfield work begins without legacy constraints, while brownfield work changes an existing system whose behaviour, interfaces and operational history may be only partly documented. The studies do not support a single greenfield-versus-brownfield multiplier.\n\nThe mechanism changes with the degree of automation. At Level 1, repository familiarity determines whether suggestions save search and typing or add context and verification overhead. At Level 2, specifications and executable acceptance checks can make structured brownfield changes unusually tractable. At Level 3, the strongest cases come from greenfield systems or repetitive migrations whose environments, tests and work units were redesigned for autonomous execution.\n\nBrownfield risk is not limited to code complexity. A controlled study of ambiguous business terminology found that an agent could preserve technical consistency while implementing the wrong domain behaviour ([ST-75](http://urn.kb.se/resolve?urn=urn:nbn:se:kau:diva-110473)). Repository context is therefore sufficient only when it contains the relevant intent; undocumented history and organisation-specific language remain human-owned inputs.\n\n## Compare the greenfield and brownfield evidence\n\n| Project context | Representative evidence | Observed effect | Interpretation |\n|---|---|---|---|\n| Bounded greenfield task | Controlled ChatGPT coding experiment | Helped on coding puzzles; no efficiency or quality gain on typical development work | Task realism changes the measured effect even within one experiment |\n| Production greenfield system | OpenAI agent-built product | One-tenth of the estimated manual development time—equivalent to 10× potential throughput | Large system claim, but based on a counterfactual rather than a control group |\n| Unfamiliar brownfield code | Two controlled student studies | 34.9% less time in one study; more tests passed in both | Implementation improves without corresponding evidence of deeper code comprehension |\n| Familiar, mature brownfield code | METR experienced-maintainer trial | 19% slower | Tacit knowledge and verification overhead can outweigh generation speed |\n| Structured brownfield change | CodeFuse, PicoScenes and migration cases | Moderate gains to multi-fold acceleration | Existing code can be highly tractable when changes repeat and tests provide an oracle |\n| Repository-wide agent adoption | Cursor matched longitudinal study | Velocity increase was large but transient | Early output gains may be offset by persistent complexity and review costs |\n\nTwo small controlled brownfield experiments make the distinction clearer. Ten undergraduates modifying an unfamiliar legacy web application completed tasks in [34.9% less time and made 50% more solution progress](https://arxiv.org/abs/2506.10051) with Copilot. A replication with 15 graduate students also found significantly lower task time and more tests passed, but [no improvement in code comprehension](https://arxiv.org/abs/2511.02922). These results differ from METR’s slowdown among expert maintainers of large repositories they knew well.\n\n**Review finding:** lifecycle status alone is a poor predictor. The strongest positive brownfield results occur when the task is deterministic, the relevant context can be supplied and correctness is executable. Brownfield work becomes less favourable when success depends on tacit architecture knowledge, broad side effects or expensive human verification. **Basis:** `E-04`\n\n, `E-18`\n\n–`E-22`\n\n, `E-39`\n\n–`E-42`\n\n.\n\n### Programming language is a moderator, but “strongly typed is faster” is not yet established\n\nThe language hypothesis is plausible: a compiler or type checker gives an agent a cheap, deterministic signal before a human reads the patch. [Google’s production completion system](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/) supports the mechanism. In Go, approximately 8% of raw suggestions failed to compile; semantic checks filtered 80% of those failures, and acceptance grew by 1.9× after the checks were introduced, compared with 1.3× in languages without them. This is evidence that compiler-like feedback improves an AI workflow—not that static typing alone increases end-to-end developer productivity.\n\nIts role grows with automation. At Level 1, compiler and linter feedback filters suggestions before the developer accepts them. At Level 2, those checks become executable acceptance constraints for a multi-file change. At Level 3, they form part of the agent’s build–test–repair loop and reduce how often a human must inspect obviously invalid states.\n\nCross-language evaluations show genuine variation, but not a simple static-versus-dynamic ranking. An evaluation across [19 programming languages](https://arxiv.org/abs/2501.02338) found above-average results for several statically typed, lower-abstraction languages. A study of all 2,033 LeetCode problems found substantial correctness differences across [C, Java, JavaScript and Python](https://doi.org/10.1145/3715108). The newer [SWE-bench Multilingual](https://www.swebench.com/multilingual.html), [Multi-SWE-bench](https://arxiv.org/abs/2504.02605) and [OmniCode](https://arxiv.org/abs/2602.02262) evaluations likewise show materially different resolution rates by language and task. In OmniCode, for example, SWE-agent’s best Java test-generation result was only 20.9%.\n\nLanguage familiarity and training-data density can move in the opposite direction to type-system strength. A controlled cross-language trajectory study found large differences in token consumption across [Python, Java, Rust and OCaml](https://arxiv.org/abs/2607.22807), with agents repeatedly producing non-compiling solutions in less familiar languages. A Java unit-test study found that type-rich context helped, but generated tests still suffered from compilation and oracle errors ([EASE 2024 evaluation](https://doi.org/10.1145/3661167.3661216)). [TypePilot](https://aclanthology.org/2025.ommm-1.11/) provides more direct evidence that a Scala type system can constrain insecure generations, although it evaluates generated-code robustness rather than developer time.\n\n**Review finding:** treat the language as part of the verification harness. Static types, exhaustive compilers and mature linters should reduce the cost of rejecting a bad change, while popular ecosystems may improve generation quality through better training coverage. No identified field experiment isolates static typing as a causal productivity multiplier.\n\n### Developer experience changes where AI helps\n\n“Experience” is not one variable. General coding tenure, familiarity with a particular repository and knowledge of the problem domain can produce different effects. In assisted-development trials, less-experienced developers often gain more. The three randomised field experiments covering 4,867 developers found [higher adoption and larger productivity gains among less-experienced developers](https://pubsonline.informs.org/doi/10.1287/mnsc.2025.00535). In the CodeFuse field experiment, the measured output increase was [concentrated among junior programmers](https://www.bis.org/publ/work1208.pdf), despite minimal junior–senior differences in suggestion acceptance.\n\nA phased enterprise rollout provides a useful counterexample. Daily usage, commit and HR records for 743 developers showed larger contribution gains among senior developers, plus spillovers to junior developers working with senior adopters ([Ye et al.](https://aisel.aisnet.org/icis2025/impl_adopt/impl_adopt/2/)). The outcome is commits rather than releases, and adoption was not random, but it reinforces the distinction between general tenure and the expertise needed to steer and verify AI-assisted work.\n\nRepository familiarity can reverse that pattern. METR’s experienced maintainers were 19% slower on repositories they had worked in for an average of five years, while two small studies of developers modifying unfamiliar brownfield applications found faster completion and more tests passed. This does not show that novices outperform experts. It suggests that assistance is more valuable when the tool supplies missing context, but can impose prompting and verification overhead when a developer already holds a strong mental model of the code.\n\n## Read the extended developer-experience evidence\n\nTwo journal studies sharpen the distinction between expertise and task context. In an online experiment with 173 programmers, code comments increased adoption of AI-generated JavaScript for both novices and experts, while expertise itself did not significantly change adoption ([ST-112](https://doi.org/10.1016/j.jss.2025.112634)). By contrast, an observational study of 686 developer–ChatGPT conversations found that larger projects and more experienced developers were more likely to obtain useful issue-resolution help, especially on simpler, well-scoped problems ([ST-107](https://doi.org/10.1007/s10664-025-10745-8)). Adoption and successful use are therefore different outcomes.\n\nAt higher degrees of automation, domain expertise becomes more—not less—important. In Anthropic’s analysis of approximately 400,000 Claude Code sessions, novice-rated sessions reached verified success 15% of the time, compared with 28–33% for intermediate-to-expert sessions. When sessions encountered trouble, verified recovery rose from 4% for novices to 15% for experts. The authors found that [task-specific expertise predicted success more strongly than working in a software occupation](https://www.anthropic.com/research/claude-code-expertise), although the outcomes were inferred from transcripts and telemetry rather than production value.\n\nThe relationship is therefore not a simple equalising effect. A study of more than 160,000 open-source developers found that observed contribution gains were [concentrated among experienced developers](https://doi.org/10.1126/science.adz9311), with no significant gain among early-career developers. Different tasks, experience measures and workflows explain part of the apparent contradiction.\n\nQualitative and longitudinal studies describe the same shift from implementation toward supervision. Experienced professionals retain control over architecture and quality rather than relying on unconstrained “vibe coding” ([Huang et al.](https://arxiv.org/abs/2512.14012)). A six-month panel found 84% reporting productivity improvement at both waves, while worsened developer experience nearly doubled in the matched cohort ([Vella and Blincoe](https://arxiv.org/abs/2605.23135)). Microsoft’s [SPACE study of more than 500 developers](https://www.microsoft.com/en-us/research/publication/the-space-of-ai-real-world-lessons-on-ais-impact-on-developers/) found benefits concentrated in routine work and weaker evidence for collaboration, while the peer-reviewed analysis of more than 200,000 repositories found self-admitted use extending well beyond code generation ([ST-60](https://doi.org/10.1109/TSE.2026.3681886)).\n\nLearning evidence adds a reason to preserve human scrutiny. A comparison of human pair programming and Copilot found similar frequencies of successful knowledge-transfer episodes, but developers examined Copilot suggestions less critically than suggestions from human partners ([ST-80](https://arxiv.org/abs/2506.04785)). In a project-based software-engineering course, reliance on AI was not correlated with grades; requirements, design, testing and code review became more—not less—important ([ST-83](https://doi.org/10.1145/3797092)). These studies do not estimate professional output, but support measuring understanding and verification alongside speed.\n\nA task-based observational study of 92 students, interns and junior developers reached a compatible conclusion. Prompt refinement, suggestion acceptance, testing and other collaboration behaviours were associated with programming outcomes, reinforcing the importance of evaluating and refining AI output rather than measuring tool access alone ([ST-77](https://lutpub.lut.fi/handle/10024/172536)).\n\nTeam effects also extend beyond individual output. A mixed-method study of 30 industrial developers, followed by a survey of another 131 professionals, found that GenAI reduced some low-level interruptions but shifted human conversations toward context, joint reasoning and social connection ([ST-109](https://doi.org/10.1109/TSE.2026.3655626)). That is a change in collaboration structure, not evidence that fewer human interactions are automatically more productive.\n\n**Review finding:** less-experienced developers may gain more from Level 1 assistance on bounded tasks, while experts may gain little—or slow down—when assistance disrupts work in a familiar repository. At Levels 2 and 3, domain knowledge improves specification, steering, exception handling and verification. Experience should therefore be measured as **coding tenure, repository familiarity and domain expertise**, not a single seniority label.\n\n## What happens to quality, security and review?\n\nQuality is a downstream productivity outcome, not a separate concern. Faster generation creates value only if review, security and maintenance costs remain controlled.\n\n### Automated review is necessary—but not a 4× cause\n\nReview responsibility changes with the degree of automation. At Level 1, the developer inspects suggestions while implementing the change. At Level 2, a human reviews and owns a completed agent-authored change. At Level 3, automated review becomes a first-pass control before accountable human approval. The third-level increase cannot be attributed to automated review alone, and the available evidence does not isolate review as a causal factor.\n\nOpenAI’s guide to AI-native engineering says automated review gives engineers more confidence that major bugs are caught, but [does not necessarily make the pull-request process faster](https://cdn.openai.com/business-guides-and-resources/building-an-ai-native-engineering-team.pdf). Humans still need to understand the implications and own the merge.\n\n## Read the extended automated-review evidence\n\nA 2026 study of 31,073 CodeRabbit review comments across 10,191 pull requests found that only [36.4% were accepted](https://arxiv.org/abs/2607.03316), while 56.3% were rejected. Common reasons included false positives, redundant comments and suggestions outside the intended scope.\n\nAtlassian reports a more favourable result from a tightly integrated review system. Its year-long evaluation across more than 1,900 repositories found that RovoDev reduced [pull-request cycle time by 30.8%](https://arxiv.org/abs/2601.01129) and human-written review comments by 35.6%; 38.7% of its comments triggered a subsequent code change. The paper was accepted to the ICSE 2026 Software Engineering in Practice track, but the evaluation was observational and conducted by the tool’s developer.\n\nThe review bottleneck is real. A GitLab survey of 1,528 developers and technology buyers found that [85% agreed AI had shifted the bottleneck](https://about.gitlab.com/press/releases/2026-06-23-gitlab-research-reveals-organizations-are-generating-ai-code-faster-than-they-can-control-it/) from writing code to review and validation. But survey agreement is not proof that an automated reviewer removes it.\n\nLarge-scale repository studies show why review results depend on task and governance. A study of [40,214 human- and agent-authored pull requests](https://doi.org/10.1145/3805760.3814909) found materially different collaboration patterns when agents authored the change. A separate analysis of [7,156 agent-authored pull requests](https://arxiv.org/abs/2602.08915) found that acceptance varied more by task type than by agent: documentation reached 82.1%, compared with 66.1% for new features. Early-adoption research covering [25,264 agentic pull requests](https://arxiv.org/abs/2607.14037) found that most repositories still produced only one or two agent-authored pull requests over three months and relied predominantly on a single human reviewer.\n\nThe updated evidence extends this pattern. Across 12,433 agent-authored pull requests, functional failures—especially specification mismatch and logic defects—were the dominant visible rejection reason, while 84.2% of rejected pull requests closed without inline feedback ([Coding Agents in the Wild](https://doi.org/10.1109/access.2026.3696573)). In a separate study of 567 Claude Code pull requests, 83.8% were merged, but 45.1% of merged changes required human revision ([On the Use of Agentic Coding](https://doi.org/10.1145/3798166)). Acceptance therefore signals usefulness, not zero review cost.\n\nIntegration remains a separate constraint. Deterministic merge simulation found conflicts in 27.67% of 107,000 processed agent pull requests ([AgenticFlict](https://doi.org/10.5281/zenodo.20118379)). In a smaller enterprise-codebase experiment, four AI review configurations produced 120 T-SQL reports and identified issues missed in the original human review, although the thesis design did not estimate end-to-end review time ([AI-generated T-SQL review study](https://www.doria.fi/handle/10024/194633)).\n\nThe broader repository record confirms that acceptance is not equivalent to wholesale integration. [AIDev](https://arxiv.org/abs/2602.09185) catalogues 932,791 agent-authored pull requests across 116,211 repositories. A study of 9,427 agentic pull requests found that core developers more consistently required passing CI before acceptance ([Cynthia et al.](https://arxiv.org/abs/2601.20106)), while a longitudinal comparison tracks changing mergeability, complexity and defect proneness in agent- and human-authored pull requests ([Morovati et al.](https://arxiv.org/abs/2607.21832)). PatchTrack found a median of only [25% of suggested code integrated](https://link.springer.com/article/10.1007/s10664-026-10869-5) across 338 pull requests; developers usually extracted, adapted or rejected parts of a patch. These studies do not estimate causal speed, but show why generated or accepted code volume can overstate delivered value.\n\nCode review also remains a social control. Interviews and simulated review sessions found that peer review shifted responsibility from individual to collective accountability, while [LLM-assisted review disrupted that reciprocal process](https://doi.org/10.1145/3721127). A separate qualitative study with 20 engineers found that LLM feedback reduced some emotional and self-regulation costs but shifted attention towards cognitive processing; the authors still argued for [human accountability and the social role of peer review](https://doi.org/10.1145/3830405). These studies do not measure review speed, but explain why automated feedback cannot by itself own a release decision.\n\nGovernance evidence points in the same direction. A mixed repository-and-survey study found that developers who declare AI-generated code often do so to support later review and debugging, while substantial modification was a common reason not to declare it ([ST-101](https://doi.org/10.1145/3771937)). In a controlled crowdsourced-testing study, agent-generated feedback improved revised reports and later first submissions, but did not remove specificity and execution friction ([ST-100](https://doi.org/10.1145/3828168)).\n\nInterventions at narrower review stages show both useful automation and displacement. Across 18,256 pull requests, AI-generated descriptions were associated with [shorter review time and a higher merge probability](https://arxiv.org/abs/2402.08967). At Google, ML-generated edits now resolve [7.5% of reviewer comments](https://research.google/pubs/resolving-code-review-comments-with-machine-learning/), an effect estimated to save hundreds of thousands of engineer hours annually at that scale. AutoCommenter was deployed across four languages to tens of thousands of Google developers and produced a measurable workflow benefit, although the published study does not translate it into an end-to-end speed multiplier ([industrial evaluation](https://research.google/pubs/ai-assisted-assessment-of-coding-practices-in-industrial-code-review/)).\n\nNegative and null results are equally instructive. In Beko’s ten-project deployment, developers resolved 73.8% of automated comments, but median pull-request closure increased from 5 hours 52 minutes to 8 hours 20 minutes ([Automated Code Review in Practice](https://arxiv.org/abs/2412.18531)). Meta’s review-comment patch system initially made reviewers more than 5% slower; moving suggestions from reviewers to authors removed that regression, and the final system achieved a 19.7% actionable-to-applied rate ([randomised safety trials and production experiment](https://arxiv.org/abs/2507.13499)).\n\nThe high-performing cases combine review with deterministic tests, static analysis, isolated environments, repository rules, short feedback loops and human ownership. Review is one control in a harness; [How to Sandbox AI Agent Code](/blog/microvm-sandbox-options-firecracker-opensandbox-smolvm-nono/) compares practical isolation boundaries for the execution layer.\n\n**Review finding:** automated review is a scaling control whose importance increases with the degree of automation—not a demonstrated cause of a 4× gain. At Level 1, narrow review automation can remove routine work but can also interrupt or slow the developer. At Level 2, it should reject deterministic defects before accountable human review. At Level 3, it is necessary to prevent generated volume overwhelming review capacity, but current systems still reject or fail to action many comments and do not remove the need for human approval. No identified study isolates automated review as an end-to-end productivity multiplier. **Basis:** `ST-41`\n\n–`ST-47`\n\n, `ST-51`\n\n–`ST-52`\n\n, `ST-61`\n\n, `ST-63`\n\n, `ST-66`\n\n, `ST-72`\n\n–`ST-73`\n\n, `ST-78`\n\n, `ST-80`\n\n–`ST-85`\n\n, `ST-91`\n\n, `ST-96`\n\n–`ST-101`\n\n.\n\n### Quality effects are mixed and may emerge later\n\nSpeed and throughput do not fully describe the productivity effect, and the quality risk changes with the degree of automation. At Level 1, a developer can reject a bad suggestion before it becomes a change. At Level 2, larger agent-authored patches increase the amount of behaviour a human must verify. At Level 3, generation and initial review both scale, making independent tests, security controls, monitoring and ownership part of the operating model rather than optional safeguards. A GitHub randomised controlled trial with 202 experienced developers found that Copilot users were [53.2% more likely to pass all ten unit tests](https://github.blog/news-insights/research/does-github-copilot-improve-code-quality-heres-what-the-data-says/). Blind review also found small but statistically significant improvements in readability, reliability, maintainability and conciseness.\n\nLongitudinal evidence is less uniformly positive. A matched difference-in-differences study of 806 repositories found that adopting Cursor produced a [large but transient increase in development velocity](https://arxiv.org/abs/2511.04427), while static-analysis warnings and code complexity increased persistently. Another open-source study found that, after Copilot adoption, core developers reviewed [6.5% more code while producing 19% less original code](https://arxiv.org/abs/2510.10165), consistent with maintenance work being redistributed rather than removed.\n\nA controlled maintenance experiment provides a useful boundary on that interpretation. When 75 participants evolved code written by someone else, the speed advantage for maintaining AI-assisted code was [small and statistically unreliable](https://link.springer.com/article/10.1007/s10664-026-10889-1). A separate analysis of approximately 302,600 verified AI-authored commits across 6,299 repositories found that [more than 15% introduced at least one detectable issue](https://arxiv.org/abs/2603.28592), and 22.7% of the tracked issues remained at the latest repository revision.\n\nSecurity evaluations consistently reject the assumption that compilable means safe. Copilot produced vulnerable output in approximately 40% of 1,689 programs in the peer-reviewed [“Asleep at the Keyboard” study](https://doi.org/10.1145/3610721). A 2024 replication found a lower but still substantial [27.25% vulnerable share](https://doi.org/10.1109/SANER60148.2024.00051). In real repositories, security weaknesses affected 29.5% of observed Python snippets and 24.2% of JavaScript snippets ([Fu et al.](https://doi.org/10.1145/3716848)), while a 44-developer controlled study found no significant improvement in secure API use ([Mousavi et al.](https://arxiv.org/abs/2607.11348)).\n\n## Read the extended quality, security and maintainability evidence\n\nEarlier journal evidence establishes the same boundary. When prompted with scenarios drawn from historical C/C++ vulnerabilities, Copilot reproduced the vulnerable code in about 33% of cases and the fixed code in 25% ([ST-115](https://doi.org/10.1007/s10664-023-10380-1)). A separate benchmark found Copilot’s correct-solution rate below the human comparison set, although its buggy solutions were generally easier to repair; the authors concluded that the tool could assist experts but become a liability when users could not identify faulty suggestions ([ST-116](https://doi.org/10.1016/j.jss.2023.111734)). These are code-quality benchmarks, not estimates of workplace defect rates.\n\nNewer comparisons do not support a simple quality penalty. Across 984 HumanEval samples, GPT-4 with advanced prompts outperformed the human reference code on several static quality metrics ([ST-71](https://doi.org/10.1109/msr66628.2025.00081)). Yet Copilot produced functionally correct code in 97.03% of another benchmark while 30.6% of snippets still contained a vulnerability ([Artificially Insecure](https://doi.org/10.1109/svcc65277.2025.11133629)). A study of more than 500,000 Python and Java samples found distinct defect profiles and more high-risk vulnerabilities in AI-generated code ([Human-Written vs. AI-Generated Code](https://arxiv.org/abs/2508.21634)). An analysis of more than 20,000 GitHub issues likewise found new vulnerabilities in standalone-LLM and agent-generated patches, with greater autonomy associated with additional risk ([How Safe Are AI-Generated Patches?](https://arxiv.org/abs/2507.02976)).\n\nThe direct journal search adds two useful boundaries. Twelve students using ChatGPT for unit-test engineering adopted different prompting strategies, but those strategies did not significantly change mutation score or test-smell measures ([ST-106](https://doi.org/10.1007/s10664-026-10898-0)). A separate 180-problem benchmark found ChatGPT 3.5 more reliable on easy and medium tasks than hard ones, while its 99-programmer survey linked perceived productivity to trust and tool use rather than measured delivery ([ST-114](https://doi.org/10.1016/j.scico.2024.103111)).\n\nRisk can accumulate across revisions rather than appearing in the first generation. SecureIterate analysed 1,250 generated software versions and reported statistically significant vulnerability growth during repeated feature, debugging, optimisation and refactoring cycles ([ST-92](https://doi.org/10.5281/zenodo.20550244)). This preprint is benchmark evidence rather than production-incident data, but it supports continuous security validation throughout an agent loop instead of a one-off check at merge.\n\nThe surrounding risks create verification work. Across 16 coding models, package hallucinations averaged [19.6%](https://www.usenix.org/publications/loginonline/we-have-package-you-comprehensive-analysis-package-hallucinations-code). A 2024 USENIX evaluation found that LLMs could assist code analysis but remained unreliable across tasks ([Fang et al.](https://www.usenix.org/conference/usenixsecurity24/presentation/fang)). Newer security-by-default work improved secure generation by up to 2.5× across six models and four security-critical languages, but this is a model-level benchmark rather than evidence that deployed teams experience fewer incidents ([GoodVibe](https://www.usenix.org/conference/usenixsecurity26/presentation/thang)). The strongest conclusion is therefore not that AI necessarily increases production bugs, but that generated code retains material defect and vulnerability risk and requires an explicit verification budget.\n\nAI can also strengthen verification when it is constrained by concrete vulnerability evidence. A peer-reviewed security-testing study generated successful exploit demonstrations for vulnerable dependencies in 49 applications, produced 24 pieces of vulnerability evidence and led to four assigned CVEs ([ST-110](https://doi.org/10.1109/TSE.2025.3632765)). This shows a useful security-testing capability, not that unconstrained generated application code is secure.\n\nAutomated tests can recover part of this cost, but generated tests need their own oracle. TestPilot generated JavaScript tests with median statement coverage of [70.2%](https://doi.org/10.1109/TSE.2023.3334955) in its 2024 journal publication. ChatTESTER increased compilable tests by 34.3% and tests with correct assertions by 18.7% relative to direct ChatGPT generation ([FSE 2024 study](https://doi.org/10.1145/3660783)). A broader evaluation across 17 Java projects found that prompt content and model choice materially changed outcomes ([Yang et al.](https://arxiv.org/abs/2406.18181)). The productive mechanism is therefore the executable feedback loop, not test text generation alone.\n\nRepository telemetry indicates that these costs may emerge slowly. GitClear’s analysis of [153 million changed lines](https://gitclear-public.s3.us-west-2.amazonaws.com/Coding-on-Copilot-2024-Developer-Research.pdf) found rising churn and copy-paste activity as AI adoption grew, although it could not identify AI authorship causally. DORA’s 2024 cross-organisation analysis associated a 25% increase in AI adoption with better documentation, code quality and review speed, but also [1.5% lower delivery throughput and 7.2% lower stability](https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report). These are associations, but they illustrate why local speed and delivery performance must be measured separately.\n\n[Google’s 2025 DORA research](https://dora.dev/research/2025/dora-report/) develops that finding: AI appears to amplify existing organisational strengths and weaknesses, with higher adoption associated with both greater throughput and greater instability. The appropriate comparison is therefore not simply with and without AI, but the same degree of automation under stronger and weaker delivery systems.\n\nTwo additional syntheses support that caution. A 49-study review found that LLM-generated refactoring remained unreliable despite promising results across code-quality tasks ([Alomari et al.](https://doi.org/10.1016/j.infsof.2025.107960)). A multivocal review of 24 potentially usable software-quality solutions judged most to be immature prototypes with a persistent gap between benchmark performance and production readiness ([Alenezi et al.](https://doi.org/10.1007/s11219-026-09754-7)).\n\nCollectively, these studies indicate that short-task quality can improve while repository-level complexity, persistent defects and expert review load still rise over time. This distinction is especially consequential in brownfield systems, where each accepted change becomes part of the next task’s context. [AI-Assisted Web Development: Where It Helps and Where It Fails](/blog/the-rise-of-ai-in-web-development/) applies the same boundary to pattern-heavy work that is easy to verify and judgement-heavy work where plausible output can hide defects.\n\n**Review finding:** faster generation does not inevitably reduce quality, but the evidence does not support treating quality as constant across degrees of automation. Level 1 trials can improve immediate correctness on bounded tasks while longitudinal studies still detect complexity, maintenance and security costs. At Levels 2 and 3, larger agent-authored changes expand the verification surface; maintaining quality requires independent tests, security controls, monitoring and accountable ownership to scale with generation. No identified study demonstrates that quality remains stable at the article’s 2× or 4.5× implementation scenarios. **Basis:** `ST-22`\n\n, `ST-27`\n\n, `ST-48`\n\n–`ST-50`\n\n, `ST-53`\n\n–`ST-55`\n\n, `ST-62`\n\n, `ST-64`\n\n–`ST-68`\n\n, `ST-71`\n\n, `ST-74`\n\n–`ST-76`\n\n, `ST-79`\n\n, `ST-87`\n\n–`ST-92`\n\n, `ST-94`\n\n–`ST-95`\n\n, `ST-100`\n\n–`ST-116`\n\n.\n\n## Productivity effect register\n\nThe complete 46-estimate register is retained for traceability. Open it to compare study design, effect, evidence weight and limitation.\n\n## Show all 46 productivity and delivery estimates\n\n| IDs | Study and outcome | Design | Effect | Evidence weight | Main limitation |\n|---|---|---|---|---|---|\n`ST-01` `E-01` |\n|\n\n`ST-02`\n\n`E-02`\n\n[Microsoft/Accenture/F100](https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2025.00535)— completed tasks`ST-03`\n\n`E-03`\n\n[Rocks Coding, Not Development](https://arxiv.org/abs/2402.05650)— task performance`ST-04`\n\n`E-04`\n\n[METR](https://arxiv.org/abs/2507.09089)— mature-repository task time`ST-05`\n\n`E-05`\n\n[PACIS assistant–agent trial](https://aisel.aisnet.org/pacis2026/ai_fow/ai_fow/13/)— completion time`ST-06`\n\n`E-06`\n\n[ANZ Bank](https://arxiv.org/abs/2402.05636)— Python-task completion`ST-07`\n\n`E-07`\n\n[Weber et al.](https://doi.org/10.1145/3661145)— requirements per minute`ST-08`\n\n`E-08`\n\n[More Code, Less Understanding](https://doi.org/10.1109/TSE.2026.3679627)— task completeness`ST-09`\n\n`E-09`\n\n[GILT](https://research.google/pubs/using-an-llm-to-help-with-code-understanding/)— code-understanding progress`ST-10`\n\n`E-10`\n\n[ClueBot](https://doi.org/10.1145/3702163.3702168)— JavaScript time and errors`ST-11`\n\n`E-11`\n\n[Writing Code vs. Shipping Code](https://www.nber.org/papers/w35275)— commits, releases and use`ST-12`\n\n`E-12`\n\n[Global diffusion and impact](https://doi.org/10.1126/science.adz9311)— quarterly contributions`ST-13`\n\n`E-13`\n\n[Generative AI and the Nature of Work](https://mackinstitute.wharton.upenn.edu/wp-content/uploads/2025/04/WTIC.2025_Nagle-Frank_Generative-AI.pdf)— work allocation`ST-14`\n\n`E-14`\n\n[Collaborative open source](https://arxiv.org/abs/2410.02091)— project output and coordination`ST-15`\n\n`E-15`\n\n[Copilot X](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5100609)— commits and repositories`ST-16`\n\n`E-16`\n\n[Workplace AI-coder deployment](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6515379)— productivity index`ST-17`\n\n`E-17`\n\n[Python–R natural experiment](https://arxiv.org/abs/2409.08379)— releases and commits`ST-18`\n\n`E-18`\n\n[Ant Group CodeFuse](https://www.bis.org/publ/work1208.pdf)— code and workflow`ST-19`\n\n`E-19`\n\n[Brownfield undergraduate study](https://arxiv.org/abs/2506.10051)— task time`ST-20`\n\n`E-20`\n\n[Brownfield replication](https://arxiv.org/abs/2511.02922)— time and tests`ST-21`\n\n`E-21`\n\n[PicoScenes](https://doi.org/10.1145/3786583.3786872)— implementation time`ST-22`\n\n`E-22`\n\n[Cursor repositories](https://arxiv.org/abs/2511.04427)— velocity`ST-23`\n\n`E-23`\n\n[Meta multi-line completion](https://arxiv.org/abs/2402.04141)— keystrokes saved`ST-24`\n\n`E-24`\n\n[Uplevel](https://uplevelteam.com/blog/genai-developers)— PR cycle and throughput`ST-25`\n\n`E-25`\n\n[NAV](https://arxiv.org/abs/2509.20353)— commit activity`ST-26`\n\n`E-26`\n\n[Three agile teams](https://arxiv.org/abs/2602.13766)— story points`ST-27`\n\n`E-27`\n\n[Echoes of AI](https://link.springer.com/article/10.1007/s10664-026-10889-1)— task time`ST-28`\n\n`E-28`\n\n[GovTech Singapore](https://arxiv.org/abs/2409.17434)— task speed`ST-29`\n\n`E-29`\n\n[Salesforce agentic SDLC](https://www.salesforce.com/news/stories/how-engineering-became-agentic/)— effective output`ST-30`\n\n`E-30`\n\n[Anthropic adoption](https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic)— merged PRs`ST-31`\n\n`E-31`\n\n[AWS Amazon Stores pilots](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/)— productivity gain`ST-32`\n\n`E-32`\n\n[UK Government](https://www.gov.uk/government/publications/ai-coding-assistant-trial/ai-coding-assistant-trial-uk-public-sector-findings-report)— reported time saved`ST-33`\n\n`E-33`\n\n[Meta CodeCompose](https://doi.org/10.1145/3643774)— accepted-code share`ST-34`\n\n`E-34`\n\n[Faros 2025](https://www.faros.ai/blog/ai-software-engineering)— tasks, PRs and review`ST-35`\n\n`E-35`\n\n[Faros 2026](https://www.faros.ai/research/ai-acceleration-whiplash)— delivery telemetry`ST-36`\n\n`E-36`\n\n[DX Q4 2025](https://getdx.com/report/ai-assisted-engineering-q4-impact-report/)— time and PR output`ST-37`\n\n`E-37`\n\n[Jellyfish](https://jellyfish.co/blog/are-ai-coding-tools-making-companies-more-productive/)— epics and effort`ST-31`\n\n`E-38`\n\n[AWS Prime Video](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/)— forecast duration`ST-38`\n\n`E-39`\n\n[Cisco](https://www.sonarsource.com/blog/ai-first-engineering-cisco/)— tech-debt workflow`ST-39`\n\n`E-40`\n\n[OpenAI](https://openai.com/index/harness-engineering/)— estimated project time`ST-29`\n\n`E-41`\n\n[Salesforce migration](https://www.salesforce.com/news/stories/how-engineering-became-agentic/)— task duration`ST-40`\n\n`E-42`\n\n[Amazon Q migration](https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone/)— estimated effort`ST-69`\n\n`E-43`\n\n[Three AI web-app tools](http://www.theseus.fi/handle/10024/923044)— development time`ST-70`\n\n`E-44`\n\n[Non-programmers with AI](https://doi.org/10.1007/s10664-026-10813-7)— task time`ST-86`\n\n`E-45`\n\n[Three in-IDE assistants](https://arxiv.org/abs/2404.12000)— task time`ST-93`\n\n`E-46`\n\n[Enterprise phased rollout](https://aisel.aisnet.org/icis2025/impl_adopt/impl_adopt/2/)— commits## How should organisations use and measure these findings?\n\nFor business planning, use the **10–30% release-output range** as the defensible base case and treat **2×** spec-driven and **4.5×** agent-native implementation as conditional upside scenarios. Budgets should include model and platform costs, workflow integration, reviewer capacity, testing, security and expected rework—not licences alone.\n\nThe literature would benefit from staged comparisons of representative work rather than further surveys of perceived speed. A robust organisational study would:\n\n- Select repeated task classes with sufficient history for a baseline.\n- Measure accepted intent through production operation, rather than coding time alone.\n- Separate active human time, elapsed time and agent runtime.\n- Include review rounds, escaped defects, rollback and rework.\n- Compare similar tasks at assisted, spec-driven and agent-native levels.\n- Report the model, repository characteristics, team composition, coding tenure, repository familiarity and domain expertise.\n\nRelevant outcome measures include lead time, work items deployed, change failure rate, reviewer minutes, post-merge rework and production incidents. Commits can diagnose activity, but are not an adequate primary productivity outcome. For a worked example of holding variables constant, exposing uncertainty and separating raw results from interpretation, see [Building the Ed-o-meter: Notes on Writing My Own LLM Benchmark](/blog/building-the-ed-o-meter-llm-eval-harness/).\n\n## What are the limitations of the AI coding productivity evidence?\n\nScreening, eligibility and formal risk-of-bias appraisal used one AI-assisted reviewer rather than two independent reviewers. The appraisal was not independently duplicated. Full reports were inspected where publicly accessible; otherwise eligibility and conservative domain judgements used the detailed publisher or repository record. Searches prioritised English-language, web-accessible studies and first-party disclosures, so relevant unpublished or inaccessible evaluations may be absent. The corpus is broad rather than exhaustively database-indexed. Publication bias is likely to be strongest in the agent-native tier, where organisations have an incentive to publish exceptional successes.\n\nThe number of independent causal estimates is much smaller than the total literature. A productivity-focused systematic review identified [39 peer-reviewed primary studies published through December 2024](https://doi.org/10.1145/3809494), using a broader productivity definition than this review’s quantitative comparator rule. A newer preregistered [meta-analysis of 14 productivity studies](https://arxiv.org/abs/2605.04779) estimated a moderate pooled effect (Hedges’ *g* = 0.33) with extreme heterogeneity (*I*2 = 99%); laboratory effects were larger than enterprise and open-source effects.\n\nThe corpus could be made larger by admitting more benchmark-only evaluations, student exercises without a human baseline, vendor surveys or narrow task studies. Those documents can explain mechanisms and boundaries, but counting them as equivalent productivity evidence would reduce rather than improve the review’s precision.\n\nThe evidence base is too heterogeneous for a pooled effect estimate. Controlled studies mostly evaluate individual assistance, while agent-native reports measure teams, workflows or individual migrations. Historical estimates can also change after a project begins, and commit or line-volume measures may reward activity rather than value. Greenfield and brownfield labels were assigned from the disclosed project setting and were not always prespecified by study authors. Experience measures also differ: job tenure, coding tenure, repository familiarity and task-specific expertise are not interchangeable. Programming-language studies predominantly measure correctness, compilation or benchmark success rather than developer time, so the language analysis is mechanistic and exploratory. The three-level autonomy taxonomy was applied for this review and has not itself been independently validated.\n\nBenchmark quality adds a separate constraint. OpenAI found contamination and test-design problems severe enough to stop using [SWE-bench Verified](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/), then estimated that approximately 30% of SWE-Bench Pro tasks were broken in a separate [agent-assisted audit and human annotation study](https://openai.com/index/separating-signal-from-noise-coding-evaluations/). Benchmark progress shows that agents can perform longer tasks under controlled conditions; it cannot be converted directly into hours saved without a human workflow comparator.\n\n## What does the evidence support?\n\n**The central finding is a conversion problem.** AI can increase implementation output substantially, but the gain shrinks across understanding, integration, review, testing, release and adoption. The largest matched study illustrates the loss: 180% more commits became 30% more releases and no increase in application usage. For end-to-end planning, **10–30% higher release output, centred on 20%**, is the most defensible observational range. It is a scenario rather than a pooled or causal forecast.\n\nLarger figures describe narrower implementation workflows. Conventional assistance has a reference value near **1.25×**; bounded spec-driven work has a conditional **2×** scenario; and agent-native delivery has a low-weight **4.5×** upside scenario, with an 18× task-specific migration at the upper end. These outcomes combine models with redesigned repositories, specifications, environments, tests, review and parallel execution. They should not be applied directly to the whole delivery system.\n\nShort, bounded tasks can become both faster and more correct, so the evidence does not support an inevitable speed-versus-quality trade-off. The recurrent risk is generation scaling faster than verification, moving work into review, integration, security and maintenance. Investment decisions should therefore measure the **verified outcome in production** and include model costs, integration, reviewer capacity, rework and incidents. As automation increases, AI coding becomes a systems-engineering problem: the limiting factor is how effectively an organisation can specify, verify, integrate and release correct work.\n\n## Evidence and document registers\n\nThe registers preserve the review chain from included study to source document. Evidence weight is summarised in the method and recorded for quantitative estimates in the productivity effect register. Supporting and contextual documents are not graded as productivity evidence.\n\n### A. Included empirical studies (116)\n\n## Show the 116 studies ordered by study ID\n\n`ST-01`\n\n—[How much does AI impact development speed?](https://arxiv.org/abs/2410.12944)— Google enterprise RCT.`ST-02`\n\n—[The Effects of Generative AI on High-Skilled Work](https://pubsonline.informs.org/doi/abs/10.1287/mnsc.2025.00535)— Microsoft, Accenture and Fortune 100 field experiments.`ST-03`\n\n—[Rocks Coding, Not Development: A Human-Centric, Experimental Evaluation of LLM-Supported SE Tasks](https://arxiv.org/abs/2402.05650).`ST-04`\n\n—[Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity](https://arxiv.org/abs/2507.09089)— METR.`ST-05`\n\n—[From Assistants to Agents: Exploring Efficiency and Human Agency in AI-Supported Programming](https://aisel.aisnet.org/pacis2026/ai_fow/ai_fow/13/).`ST-06`\n\n—[The Impact of AI Tool on Engineering at ANZ Bank](https://arxiv.org/abs/2402.05636).`ST-07`\n\n—[Significant Productivity Gains through Programming with Large Language Models](https://doi.org/10.1145/3661145).`ST-08`\n\n—[More Code, Less Understanding? On the Impact of AI Assistants on Developers’ Productivity and Code Ownership](https://doi.org/10.1109/TSE.2026.3679627).`ST-09`\n\n—[Using an LLM to Help With Code Understanding](https://research.google/pubs/using-an-llm-to-help-with-code-understanding/)— Google and Carnegie Mellon University.`ST-10`\n\n—[ClueBot: A Conversational Agent to Assist in JavaScript Programming](https://doi.org/10.1145/3702163.3702168).`ST-11`\n\n—[Writing Code vs. Shipping Code: The Impact of AI on Software Development](https://www.nber.org/papers/w35275)— matched event study of more than 100,000 developers.`ST-12`\n\n—[The Global Diffusion and Impact of Artificial Intelligence Writing Code](https://doi.org/10.1126/science.adz9311)—*Science*.`ST-13`\n\n—[Generative AI and the Nature of Work](https://mackinstitute.wharton.upenn.edu/wp-content/uploads/2025/04/WTIC.2025_Nagle-Frank_Generative-AI.pdf).`ST-14`\n\n—[The Impact of Generative AI on Collaborative Open-Source Software Development](https://arxiv.org/abs/2410.02091).`ST-15`\n\n—[The Impact of Generative AI on Software Development: Evidence from GitHub Copilot X](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5100609).`ST-16`\n\n—[The Productivity Effects of Generative AI: Evidence from a Workplace Deployment of an AI Coder](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6515379).`ST-17`\n\n—[The Impact of Generative AI on Open-Source Software Development](https://arxiv.org/abs/2409.08379).`ST-18`\n\n—[Generative AI and labour productivity: a field experiment on coding](https://www.bis.org/publ/work1208.pdf)— BIS and Ant Group CodeFuse study.`ST-19`\n\n—[The Effects of GitHub Copilot on Computing Students in Brownfield Programming Tasks](https://arxiv.org/abs/2506.10051).`ST-20`\n\n—[Comprehension–Performance Gap in GenAI-Assisted Brownfield Programming](https://arxiv.org/abs/2511.02922).`ST-21`\n\n—[Less Effort, More Productivity: Lessons Learned from Developing Millions of Lines of Code with Large Language Model](https://doi.org/10.1145/3786583.3786872)— ICSE 2026.`ST-22`\n\n—[Speed at the Cost of Quality? The Impact of LLM Agent Assistance on Software Development](https://arxiv.org/abs/2511.04427).`ST-23`\n\n—[Multi-line AI-assisted Code Authoring](https://arxiv.org/abs/2402.04141)— Meta.`ST-24`\n\n—[Does GenAI Improve Software Developer Productivity?](https://uplevelteam.com/blog/genai-developers)— Uplevel Data Labs.`ST-25`\n\n—[Longitudinal GitHub Copilot Adoption at NAV IT](https://arxiv.org/abs/2509.20353).`ST-26`\n\n—[Longitudinal Multi-case Study of AI-assisted Agile Teams](https://arxiv.org/abs/2602.13766).`ST-27`\n\n—[Echoes of AI: Investigating the downstream effects of AI assistants on software maintainability](https://link.springer.com/article/10.1007/s10664-026-10889-1).`ST-28`\n\n—[Harnessing the Potential of Gen-AI Coding Assistants in Public Sector Software Development](https://arxiv.org/abs/2409.17434)— GovTech Singapore.`ST-29`\n\n—[Pioneering the Agentic Shift Within Salesforce Engineering](https://www.salesforce.com/news/stories/how-engineering-became-agentic/).`ST-30`\n\n—[How AI Is Transforming Work at Anthropic](https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic)— survey, interviews and internal telemetry.`ST-31`\n\n—[How frontier teams are reinventing AI-native development](https://aws.amazon.com/blogs/machine-learning/how-frontier-teams-are-reinventing-ai-native-development/)— AWS.`ST-32`\n\n—[AI Coding Assistant Trial: UK Public Sector Findings](https://www.gov.uk/government/publications/ai-coding-assistant-trial/ai-coding-assistant-trial-uk-public-sector-findings-report).`ST-33`\n\n—[AI-assisted Code Authoring at Scale](https://doi.org/10.1145/3643774)— Meta CodeCompose, FSE 2024.`ST-34`\n\n—[The AI Productivity Paradox](https://www.faros.ai/blog/ai-software-engineering)— Faros 2025.`ST-35`\n\n—[AI Acceleration Whiplash](https://www.faros.ai/research/ai-acceleration-whiplash)— Faros 2026.`ST-36`\n\n—[State of AI-assisted Engineering, Q4 2025](https://getdx.com/report/ai-assisted-engineering-q4-impact-report/)— DX.`ST-37`\n\n—[Are AI Coding Tools Making Companies More Productive?](https://jellyfish.co/blog/are-ai-coding-tools-making-companies-more-productive/)— Jellyfish.`ST-38`\n\n—[AI-First Engineering: How Cisco Reached Tech Debt Zero](https://www.sonarsource.com/blog/ai-first-engineering-cisco/)— SonarSource case study.`ST-39`\n\n—[Harness engineering: leveraging Codex in an agent-first world](https://openai.com/index/harness-engineering/)— OpenAI.`ST-40`\n\n—[Amazon Q Developer just reached a $260 million milestone](https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone/)— AWS production migration disclosure.`ST-41`\n\n—[RovoDev Code Reviewer: A Large-Scale Online Evaluation](https://arxiv.org/abs/2601.01129)— Atlassian and Monash University, ICSE 2026.`ST-42`\n\n—[Is Agentic Code Review Helpful?](https://arxiv.org/abs/2607.03316)`ST-43`\n\n—[When Code Authors Are Agents: A Large-Scale Study of Human–Agent Collaboration in Pull Requests](https://doi.org/10.1145/3805760.3814909).`ST-44`\n\n—[Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance](https://arxiv.org/abs/2602.08915).`ST-45`\n\n—[Early Adoption of Agentic Coding Tools by GitHub Projects](https://arxiv.org/abs/2607.14037).`ST-46`\n\n—[AI-generated Pull-request Descriptions and Review Outcomes](https://arxiv.org/abs/2402.08967).`ST-47`\n\n—[Automated Code Review in Practice](https://arxiv.org/abs/2412.18531).`ST-48`\n\n—[Does GitHub Copilot improve code quality?](https://github.blog/news-insights/research/does-github-copilot-improve-code-quality-heres-what-the-data-says/)— GitHub randomised controlled trial.`ST-49`\n\n—[AI-assisted Programming May Decrease the Productivity of Experienced Developers by Increasing Maintenance Burden](https://arxiv.org/abs/2510.10165).`ST-50`\n\n—[Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild](https://arxiv.org/abs/2603.28592).`ST-51`\n\n—[AI-assisted Assessment of Coding Practices in Industrial Code Review](https://research.google/pubs/ai-assisted-assessment-of-coding-practices-in-industrial-code-review/)— Google AutoCommenter.`ST-52`\n\n—[AI-Assisted Fixes to Code Review Comments at Scale](https://arxiv.org/abs/2507.13499)— Meta.`ST-53`\n\n—[Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions](https://doi.org/10.1145/3610721)— Communications of the ACM.`ST-54`\n\n—[Understanding the Impact of AI Code Assistants on Security API Usage](https://arxiv.org/abs/2607.11348).`ST-55`\n\n—[GoodVibe: Security-by-Vibe for LLM-Based Code Generation](https://www.usenix.org/conference/usenixsecurity26/presentation/thang)— USENIX Security 2026.`ST-56`\n\n—[How AI assistance impacts the formation of coding skills](https://www.anthropic.com/research/AI-assistance-coding-skills)— Anthropic randomised trial.`ST-57`\n\n—[Examining the Use and Impact of an AI Code Assistant in the Enterprise](https://research.ibm.com/publications/examining-the-use-and-impact-of-an-ai-code-assistant-on-developer-productivity-and-experience-in-the-enterprise)— IBM Research, CHI 2025.`ST-58`\n\n—[Dear Diary: A Randomized Controlled Trial of Generative AI Coding Tools in the Workplace](https://www.microsoft.com/en-us/research/publication/dear-diary-a-randomized-controlled-trial-of-generative-ai-coding-tools-in-the-workplace/)— Microsoft Research, ICSE 2025.`ST-59`\n\n—[Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming](https://www.microsoft.com/en-us/research/?p=894132).`ST-60`\n\n—[Self-Admitted GenAI Usage in Open-Source Software](https://doi.org/10.1109/tse.2026.3681886).`ST-61`\n\n—[A Study of Library Usage in Agent-Authored Pull Requests](https://kclpure.kcl.ac.uk/portal/en/publications/9491a361-b84f-4d93-8b55-a7ed5cca4878).`ST-62`\n\n—[Agentic AI-Driven Developer Experience for Telecom Capabilities](http://urn.kb.se/resolve?urn=urn:nbn:se:uu:diva-590043).`ST-63`\n\n—[Agentic Much? Adoption of Coding Agents on GitHub](https://doi.org/10.1145/3822180).`ST-64`\n\n—[AI based software re-engineering: case AQUATOX](https://trepo.tuni.fi/handle/10024/236191).`ST-65`\n\n—[Leveraging Prompt Engineering with AI Coding Assistants to Develop Energy-Efficient Code](https://doi.org/10.36227/techrxiv.175339126.69681777/v1).`ST-66`\n\n—[Analyzing Developer Use of ChatGPT Generated Code in Open Source GitHub Projects](https://doi.org/10.1145/3643991.3645072).`ST-67`\n\n—[Artificially Insecure: Examining GitHub Copilot’s AI-based Vulnerability Prevention System](https://doi.org/10.1109/svcc65277.2025.11133629).`ST-68`\n\n—[Assessing GitHub Copilot in Solidity Development](https://doi.org/10.1109/access.2024.3486365).`ST-69`\n\n—[Building web applications with AI: a comparative study](http://www.theseus.fi/handle/10024/923044).`ST-70`\n\n—[Can generative AI bridge the gap?](https://doi.org/10.1007/s10664-026-10813-7)`ST-71`\n\n—[Can LLMs Generate Higher Quality Code Than Humans?](https://doi.org/10.1109/msr66628.2025.00081)`ST-72`\n\n—[Coding Agents in the Wild](https://doi.org/10.1109/access.2026.3696573).`ST-73`\n\n—[The influence of AI programming assistants on collaboration in software teams](https://resolver.obvsg.at/urn:nbn:at:at-ubl:1-104661).`ST-74`\n\n—[Disrupting Test Development with AI Assistants](https://doi.org/10.1109/iaict65714.2025.11101520).`ST-75`\n\n—[Domain ambiguity in AI-assisted software development](http://urn.kb.se/resolve?urn=urn:nbn:se:kau:diva-110473).`ST-76`\n\n—[A comparative security evaluation of manual and AI-assisted web development](http://urn.kb.se/resolve?urn=urn:nbn:se:lnu:diva-147734).`ST-77`\n\n—[Exploring and modelling human–AI collaboration effectiveness in software engineering](https://lutpub.lut.fi/handle/10024/172536).`ST-78`\n\n—[Exploring the Challenges and Opportunities of AI-assisted Codebase Generation](https://arxiv.org/abs/2508.07966).`ST-79`\n\n—[Fork-Triggerable AI Coding Agents in CI](https://github.com/cpeoples/grackle).`ST-80`\n\n—[From Developer Pairs to AI Copilots](https://arxiv.org/abs/2506.04785).`ST-81`\n\n—[From Meeting to Pull Request](http://urn.kb.se/resolve?urn=urn:nbn:se:miun:diva-57923).`ST-82`\n\n—[From Preventive to Reactive: How AI Coding Assistants Transform Developers’ Security Awareness](https://doi.org/10.13016/m2gqyg-hl1c).`ST-83`\n\n—[From Specifications to Implementation in the Gen-AI Era](https://doi.org/10.1145/3797092).`ST-84`\n\n—[GitHub Copilot Developer Experience: Novice and Experienced Developers](https://trepo.tuni.fi/handle/10024/231602).`ST-85`\n\n—[How Do Developers Interact with AI?](https://doi.org/10.1145/3808120)`ST-86`\n\n—[How far are AI-powered programming assistants from meeting developers’ needs?](https://arxiv.org/abs/2404.12000)`ST-87`\n\n—[How Safe Are AI-Generated Patches?](https://arxiv.org/abs/2507.02976)`ST-88`\n\n—[Human-Written vs. AI-Generated Code](https://arxiv.org/abs/2508.21634).`ST-89`\n\n—[LLM impact on blind and low-vision programming](https://arxiv.org/abs/2504.17018).`ST-90`\n\n—[Multi-Simulation as Labor Compression](https://doi.org/10.5281/zenodo.21437017).`ST-91`\n\n—[On the Use of Agentic Coding](https://doi.org/10.1145/3798166).`ST-92`\n\n—[Security Weaknesses in LLM-Generated Source Code](https://doi.org/10.5281/zenodo.20550244).`ST-93`\n\n—[Seniority, Spillovers, and AI-Enhanced Code Contributions](https://aisel.aisnet.org/icis2025/impl_adopt/impl_adopt/2/).`ST-94`\n\n—[Tab to Autocomplete: AI Coding Assistants and Web Accessibility](https://doi.org/10.1145/3663548.3688513).`ST-95`\n\n—[The Impact of Generative AI Coding Assistants on Developers Who Are Visually Impaired](https://doi.org/10.1145/3706598.3714008).`ST-96`\n\n—[Understanding and Predicting Accepted Code Suggestions in AI-Assisted Programming](https://doi.org/10.1145/3808125).`ST-97`\n\n—[AgenticFlict: A Large-Scale Dataset of Merge Conflicts in AI Coding Agent Pull Requests](https://doi.org/10.5281/zenodo.20118379).`ST-98`\n\n—[Evaluating AI-Generated T-SQL Code Review Reports in an Enterprise-Scale Codebase](https://www.doria.fi/handle/10024/194633).`ST-99`\n\n—[What Do AI Agents Actually Change?](https://doi.org/10.1007/978-3-032-30699-9_12)`ST-100`\n\n—[More than a Judge: Agent–Human Interaction in Crowdsourced Testing Assessment](https://doi.org/10.1145/3828168).`ST-101`\n\n—[On Developers’ Self-Declaration of AI-Generated Code](https://doi.org/10.1145/3771937).`ST-102`\n\n—[Understanding and Enhancing CS Students’ Interaction Experience with AI Coding Assistant Tools](https://doi.org/10.1145/3785479).`ST-103`\n\n—[Unveiling the Role of ChatGPT in Software Development](https://doi.org/10.1145/3798163).`ST-104`\n\n—[AI Support for Data Scientists](https://doi.org/10.1007/s10664-025-10622-4).`ST-105`\n\n—[Developers’ Shared Conversations with ChatGPT in GitHub Pull Requests and Issues](https://doi.org/10.1007/s10664-024-10540-x).`ST-106`\n\n—[How Students Use Generative AI for Software Testing](https://doi.org/10.1007/s10664-026-10898-0).`ST-107`\n\n—[What Makes ChatGPT Effective for Software Issue Resolution?](https://doi.org/10.1007/s10664-025-10745-8)`ST-108`\n\n—[AI-Assisted Collaboration: GitHub Copilot and Windsurf](https://doi.org/10.1109/MS.2026.3671677).`ST-109`\n\n—[From Disruptions to Discussions](https://doi.org/10.1109/TSE.2026.3655626).`ST-110`\n\n—[How Can ChatGPT Support Human Security Testers?](https://doi.org/10.1109/TSE.2025.3632765)`ST-111`\n\n—[LLM-Based Test-Driven Interactive Code Generation](https://doi.org/10.1109/TSE.2024.3428972).`ST-112`\n\n—[Do Comments and Expertise Still Matter?](https://doi.org/10.1016/j.jss.2025.112634)`ST-113`\n\n—[Problems, Causes and Solutions in AI Pair Programming](https://doi.org/10.1016/j.jss.2024.112204).`ST-114`\n\n—[“Will I Be Replaced?” ChatGPT and Programmer Perceptions](https://doi.org/10.1016/j.scico.2024.103111).`ST-115`\n\n—[Is GitHub’s Copilot as Bad as Humans at Introducing Vulnerabilities in Code?](https://doi.org/10.1007/s10664-023-10380-1)`ST-116`\n\n—[GitHub Copilot AI Pair Programmer: Asset or Liability?](https://doi.org/10.1016/j.jss.2023.111734)\n\n### B. Systematic reviews and meta-analyses (4)\n\n## Show the 4 systematic reviews and meta-analyses\n\n[The Impact of LLM-Assistants on Software Developer Productivity: A Systematic Review and Mapping Study](https://doi.org/10.1145/3809494).[A meta-analysis of the effect of generative AI on productivity and learning in programming](https://arxiv.org/abs/2605.04779)— preregistered (PRISMA-P, OSF, December 2025).[Using LLMs to Enhance Code Quality: A Systematic Literature Review](https://doi.org/10.1016/j.infsof.2025.107960).[Generative AI Solutions for Software Quality: Assessing Industrial Readiness](https://doi.org/10.1007/s11219-026-09754-7).\n\n### C. Supporting and contextual documents (61)\n\n## Show the 61 supporting documents alphabetically\n\n[A Large-Scale Survey on the Usability of AI Programming Assistants](https://doi.org/10.1145/3597503.3608128)— ICSE 2024.[Accountability in Code Review: The Role of Intrinsic Drivers and the Impact of LLMs](https://doi.org/10.1145/3721127).[Agentic coding and persistent returns to expertise](https://www.anthropic.com/research/claude-code-expertise)— analysis of approximately 400,000 Claude Code sessions.[AI Accountability Report](https://about.gitlab.com/press/releases/2026-06-23-gitlab-research-reveals-organizations-are-generating-ai-code-faster-than-they-can-control-it/)— GitLab and The Harris Poll.[AIDev: Studying AI Coding Agents on GitHub](https://arxiv.org/abs/2602.09185).[An Empirical Evaluation of LLMs for Automated Unit Test Generation](https://doi.org/10.1109/TSE.2023.3334955)— IEEE TSE 2024.[Anthropic Economic Index: AI’s Impact on Software Development](https://www.anthropic.com/research/impact-software-development).[Are Prompt Engineering and TODO Comments Friends or Foes?](https://lab-design.github.io/papers/ICSE-24b/).[Are We All Using Agents the Same Way?](https://arxiv.org/abs/2601.20106).[Assessing and Analyzing the Correctness of GitHub Copilot’s Code Suggestions](https://doi.org/10.1145/3715108).[Assessing the Security of GitHub Copilot Generated Code: A Targeted Replication](https://doi.org/10.1109/SANER60148.2024.00051)— SANER 2024.[Building an AI-native engineering team](https://cdn.openai.com/business-guides-and-resources/building-an-ai-native-engineering-team.pdf)— OpenAI.[Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows](https://arxiv.org/abs/2507.08149).[Coding on Copilot: 2024 Developer Research](https://gitclear-public.s3.us-west-2.amazonaws.com/Coding-on-Copilot-2024-Developer-Research.pdf)— GitClear.[Continuance Use of AI Coding Assistants Among South Korean Industry Developers](https://doi.org/10.1007/s10664-025-10730-1).[Developer Ecosystem Survey 2025](https://devecosystem-2025.jetbrains.com/)— JetBrains.[Developer Interaction Patterns with Proactive AI: A Five-Day Field Study](https://doi.org/10.1145/3742413.3789148).[DORA Accelerate State of DevOps Report 2024](https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report).[Empirical Analysis of Generative AI Tool Adoption in Software Development](https://doi.org/10.1016/j.infsof.2026.108036).[Engagement in Code Review: Peer Versus LLM Interactions](https://doi.org/10.1145/3830405).[Evaluating and Improving ChatGPT for Unit Test Generation](https://doi.org/10.1145/3660783)— FSE 2024.[Evaluation of ChatGPT-4 Code Generation Across 19 Programming Languages](https://arxiv.org/abs/2501.02338).[GitTaskBench](https://arxiv.org/abs/2508.18993).[How Do AI Coding Agents Contribute to Software Development?](https://arxiv.org/abs/2607.21832).[How Meta Used AI to Map Tribal Knowledge in Large-scale Data Pipelines](https://engineering.fb.com/2026/04/06/developer-tools/how-meta-used-ai-to-map-tribal-knowledge-in-large-scale-data-pipelines/).[How Unified AI Agents Optimise Performance at Hyperscale](https://engineering.fb.com/2026/04/16/developer-tools/capacity-efficiency-at-meta-how-unified-ai-agents-optimize-performance-at-hyperscale/)— Meta.[Introducing the SWE-Lancer Benchmark](https://openai.com/index/swe-lancer/)— OpenAI.[Investigating the Role of Cultural Values in Adopting LLMs for Software Engineering](https://doi.org/10.1145/3725529).[Large Language Models for Code Analysis: Do LLMs Really Do Their Job?](https://www.usenix.org/conference/usenixsecurity24/presentation/fang)— USENIX Security 2024.[Measuring GitHub Copilot’s Impact on Productivity](https://doi.org/10.1145/3633453).[ML-Enhanced Code Completion Improves Developer Productivity](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/)— Google, 2022.[Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving](https://arxiv.org/abs/2504.02605).[Need Help? Designing Proactive AI Assistants for Programming](https://www.microsoft.com/en-us/research/publication/need-help-designing-proactive-ai-assistants-for-programming/).[Navigating the Complexity of Generative AI Adoption in Software Engineering](https://doi.org/10.1145/3652154).[OmniCode: A Benchmark for Evaluating Software Engineering Agents](https://arxiv.org/abs/2602.02262).[On the Evaluation of Large Language Models in Unit Test Generation](https://arxiv.org/abs/2406.18181).[PatchTrack: A Comprehensive Analysis of ChatGPT’s Influence on Pull-request Outcomes](https://link.springer.com/article/10.1007/s10664-026-10869-5).[Professional Software Developers Don’t Vibe, They Control](https://arxiv.org/abs/2512.14012).[Resolving Code Review Comments with Machine Learning](https://research.google/pubs/resolving-code-review-comments-with-machine-learning/)— Google.[Security Weaknesses of Copilot-Generated Code in GitHub Projects](https://doi.org/10.1145/3716848)— ACM TOSEM 2025.[Separating Signal from Noise in Coding Evaluations](https://openai.com/index/separating-signal-from-noise-coding-evaluations/)— OpenAI.[Stack Overflow Developer Survey 2024: AI](https://survey.stackoverflow.co/2024/ai).[Stack Overflow Developer Survey 2025: AI](https://survey.stackoverflow.co/2025/ai).[State of AI-assisted Software Development 2025](https://dora.dev/research/2025/dora-report/)— DORA.[State of Developer Experience Report 2025](https://www.atlassian.com/blog/developer/developer-experience-report-2025)— Atlassian.[Still Just Personal Assistants? Generative AI Adoption in Software Organisations](https://doi.org/10.1016/j.infsof.2025.107805).[Survey: The AI Wave Continues to Grow](https://github.blog/news-insights/research/survey-ai-wave-grows/)— GitHub 2024.[SWE-agent: Agent–Computer Interfaces Enable Automated Software Engineering](https://arxiv.org/abs/2405.15793).[SWE-bench Multilingual](https://www.swebench.com/multilingual.html).[SWE-bench Multimodal](https://arxiv.org/abs/2410.03859).[SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://proceedings.iclr.cc/paper_files/paper/2024/hash/edac78c3e300629acfe6cbe9ca88fb84-Abstract-Conference.html)— ICLR 2024.[SWE-Skills-Bench](https://arxiv.org/abs/2603.15401).[The Best Programming Language for Tokenmaxxing](https://arxiv.org/abs/2607.22807).[The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study](https://arxiv.org/abs/2605.23135).[The SPACE of AI: Real-World Lessons on AI’s Impact on Developers](https://www.microsoft.com/en-us/research/publication/the-space-of-ai-real-world-lessons-on-ais-impact-on-developers/).[TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code](https://aclanthology.org/2025.ommm-1.11/).[Usage, Effects and Requirements for AI Coding Assistants in the Enterprise](https://doi.org/10.1145/3786181.3788727).[Using AI-Based Coding Assistants in Practice](https://doi.org/10.1016/j.infsof.2024.107610).[Using Large Language Models to Generate JUnit Tests](https://doi.org/10.1145/3661167.3661216)— EASE 2024.[We Have a Package for You: Package Hallucinations by Code-generating LLMs](https://www.usenix.org/publications/loginonline/we-have-package-you-comprehensive-analysis-package-hallucinations-code).[Why SWE-bench Verified No Longer Measures Frontier Coding Capabilities](https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/)— OpenAI.\n\n### D. Formal risk-of-bias register\n\nThe study-level appraisal is available as a [downloadable CSV](/research/ai-coding-productivity/risk-of-bias/risk-of-bias-register.csv). Studies contributing more than one productivity estimate retain all applicable `E`\n\nidentifiers in the same row. The accompanying [method note](/research/ai-coding-productivity/risk-of-bias/README.md) records instrument routing, overall-judgement rules, reviewer limitations and reproduction steps.\n\n### E. Machine-readable knowledge bundle\n\nThe full set of registers above — search logs, screening and eligibility decisions, risk-of-bias judgements and evidence-weight grades — is also published as an [OKF knowledge bundle](/research/ai-coding-productivity/okf/index.md), one cross-linked document per dataset with provenance and schema metadata, for agents and tools that want to traverse the evidence chain programmatically rather than parse the CSVs directly.", "url": "https://wpnews.pro/news/how-much-does-ai-improve-software-development-productivity", "canonical_source": "https://reinvently.co.uk/blog/ai-coding-productivity-evidence/", "published_at": "2026-07-28 22:31:06+00:00", "updated_at": "2026-08-23 17:43:50.191079+00:00", "lang": "en", "topics": ["artificial-intelligence", "developer-tools", "ai-research"], "entities": ["Google", "NBER", "AWS", "Salesforce", "DORA"], "alternates": {"html": "https://wpnews.pro/news/how-much-does-ai-improve-software-development-productivity", "markdown": "https://wpnews.pro/news/how-much-does-ai-improve-software-development-productivity.md", "text": "https://wpnews.pro/news/how-much-does-ai-improve-software-development-productivity.txt", "jsonld": "https://wpnews.pro/news/how-much-does-ai-improve-software-development-productivity.jsonld"}}