{"slug": "how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it", "title": "How AI Can Lead to a Decline in Code Quality and How to Fix It", "summary": "A developer argues that AI-assisted coding shifts the bottleneck from writing code to reviewing it, since generation has become cheap while understanding and verification still demand engineering time. Using a C# payment-retry example, the post shows how a plausible-looking generated loop can double-charge a customer when a gateway response times out, and proposes a stable idempotency key reused across retries as a safer design. It cites a GitHub Copilot experiment in which professional developers finished one task 55% faster on average and DORA's 2024 research linking AI adoption to reported individual productivity gains alongside negative effects on delivery stability and throughput.", "body_md": "*A practical look at review capacity, hidden edge cases, and habits that help teams maintain quality.*\n\nThe principles in this post apply across programming languages. The examples use C# because it is the language I’m most familiar with.\n\nAI-assisted coding has changed the cost of developing software. For many tasks, that is genuinely useful. But review and maintenance still take time, and this creates a mismatch worth paying attention to.\n\nConsider a hypothetical pull request for a payment retry feature. The assistant produces a large, polished change quickly. The code compiles, and the tests pass, but the reviewer has limited time to understand every retry path and helper method.\n\nWeeks later after this feature get's released to production a customer accidently get's charged twice. **How could I have missed this?**\n\nThe risk is not that AI always writes bad code. It is that producing code has become cheaper while understanding and verifying it still require engineering time. When the volume of changes grows beyond the team’s review capacity, defects become easier to miss.\n\nAI coding tools can help developers finish certain tasks faster. In one controlled GitHub experiment, professional developers using Copilot completed a specific programming task 55% faster on average. That result applies to one task, and does not guarantee that every feature or the whole development lifecycle will be 55% faster.<sup>1</sup>\n\nThe risk appears when code production speeds up but review capacity does not.\n\n```\nIllustrative, not measured data\n\nCode arriving:  ███████████████\nCareful review: ██████\n                └── The rest waits or gets less attention\n```\n\nWhen a pull request is too large to understand in the time available, reviewers can end up checking whether it *looks* right instead of working through what it actually does. Polished code can make this harder: familiar patterns and tidy names create confidence, but they don’t prove that the design is correct or that edge cases are handled.\n\nDORA’s 2024 research describes a similar tension at the delivery level. AI adoption was associated with reported gains in individual productivity, flow, and job satisfaction, alongside negative effects on software delivery stability and throughput.<sup>2</sup>\n\nThat is not proof that AI inevitably lowers code quality. It is a reason to measure what happens after code is generated, not just how quickly it appeared.\n\nImagine asking an assistant:\n\nRetry a failed payment request up to three times.\n\nIt might generate something like this:\n\n``` js\npublic async Task ChargeOrderAsync(Order order)\n{\n    for (var attempt = 0; attempt < 3; attempt++)\n    {\n        try\n        {\n            await _paymentGateway.ChargeAsync(order.Amount);\n            return;\n        }\n        catch (Exception) when (attempt < 2)\n        {\n            // Retry after a failure.\n        }\n    }\n}\n```\n\nThe loop is easy to follow. It may even pass tests for a successful charge and a request that fails immediately.\n\nBut what if the payment gateway processes the charge, then the response times out before the application receives it? The next attempt may charge the customer again.\n\nThe problem is not that the code is messy. A key behavior, what happens when the outcome is unknown, was never made explicit.\n\nA safer design usually needs a stable idempotency key for the logical order, reused across retries. The exact API depends on the payment provider, but the idea might look like this:\n\n``` js\npublic async Task ChargeOrderAsync(Order order)\n{\n    var idempotencyKey = $\"order:{order.Id}\";\n\n    for (var attempt = 0; attempt < 3; attempt++)\n    {\n        try\n        {\n            await _paymentGateway.ChargeAsync(order.Amount, idempotencyKey);\n            return;\n        }\n        catch (Exception exception)\n            when (IsRetryable(exception) && attempt < 2)\n        {\n            continue;\n        }\n    }\n}\n```\n\nThis is illustrative pseudocode, not a drop-in implementation. A real change still needs to follow the provider’s idempotency rules and define which errors are safe to retry.\n\nGoogle’s code review guidance recommends looking beyond whether code “works”: reviewers should consider design, complexity, tests, context, and whether they understand every line.<sup>3</sup> Those questions matter especially when a patch was generated quickly and looks convincing at a glance.\n\nStart with questions, not implementation:\n\nEvery AI assistant usually has a plan mode. Ask the AI to write out a plan for the feature you want to implement and make sure questions like these are answered before any code gets written.\n\nSplit work into focused changes: one for the behavior, another for an unrelated cleanup, and another for follow-up documentation.\n\nA small pull request gives a reviewer a better chance to understand how the pieces fit together. It also makes it easier to identify what caused a failure later. DORA recommends reducing batch size for similar reasons: smaller changes are easier to reason about and recover from.<sup>4</sup>\n\nFor the payment example, useful tests might cover:\n\nDon’t stop at “the tests pass.” Ask whether a test would fail if the bug you’re worried about came back. Tests are code too, and a test that repeats the implementation’s assumptions may confirm the wrong behavior.\n\nThe developer submitting a pull request should be able to explain what every changed part does, why it belongs, and what risks remain.\n\nAI can help summarize a diff or suggest review questions, but that summary should not replace reading the diff. Treat it as a map, not proof that you have visited every place on it.\n\nLines generated and tasks started are easy to count; they don’t tell you whether the change helped users or became expensive to maintain.\n\nLook at several signals together: review time, rework, changes that need to be rolled back or urgently fixed, and whether delivery is becoming more or less stable. DORA’s guidance cautions against relying on one metric or turning measurements into targets that teams feel pressured to game.<sup>4</sup>\n\nThese are delivery signals, not direct measures of code quality. Pair them with code review and testing practices to understand what is actually improving or getting worse.\n\n```\nDefine behavior\n      ↓\nAsk for a plan\n      ↓\nGenerate one small change\n      ↓\nTest edge cases and inspect the diff\n      ↓\nHuman review and ownership\n      ↓\nMerge, monitor, and learn\n```\n\nBefore approving, ask:\n\nIf the honest answer to “Do I understand this?” is no, ask for clarification or a smaller change before approving it.\n\nAI can shorten the distance between an idea and a working first draft. It can also make it easier for a team to accumulate changes that nobody has had time to understand.\n\nThe useful goal is to keep the size and arrival rate of changes within the team’s ability to review them carefully. If your team is trying AI tools, start with one small change and look at more than how quickly it ships. Check the tests, review effort, and cost of maintaining the result too.\n\nGitHub’s experiment recruited 95 professional developers to complete a JavaScript HTTP server task. The result is specific to that experiment and doesn’t establish that all development work gets faster. ↩\n\nDORA reports associations from its research; those findings don’t establish that AI alone caused changes in delivery performance. ↩\n\nGoogle’s guidance was written for code review generally, not specifically for AI-generated changes. ↩\n\nDORA presents these as delivery performance signals, not direct measures of code quality. ↩", "url": "https://wpnews.pro/news/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it", "canonical_source": "https://dev.to/mitar_nik/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it-1j0j", "published_at": "2026-09-30 09:25:10+00:00", "updated_at": "2026-09-30 09:47:58.798812+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-products"], "entities": ["GitHub", "Copilot", "DORA", "C#"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it", "markdown": "https://wpnews.pro/news/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it.md", "text": "https://wpnews.pro/news/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it.txt", "jsonld": "https://wpnews.pro/news/how-ai-can-lead-to-a-decline-in-code-quality-and-how-to-fix-it.jsonld"}}