cd /news/ai-tools/how-ai-can-lead-to-a-decline-in-code… · home › topics › ai-tools › article
[ARTICLE · art-142404] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

How AI Can Lead to a Decline in Code Quality and How to Fix It

A developer argues that AI-assisted coding shifts the bottleneck from writing code to reviewing it, since generation has become cheap while understanding and verification still demand engineering time. Using a C# payment-retry example, the post shows how a plausible-looking generated loop can double-charge a customer when a gateway response times out, and proposes a stable idempotency key reused across retries as a safer design. It cites a GitHub Copilot experiment in which professional developers finished one task 55% faster on average and DORA's 2024 research linking AI adoption to reported individual productivity gains alongside negative effects on delivery stability and throughput.

by read6 min views3 publishedSep 30, 2026

A practical look at review capacity, hidden edge cases, and habits that help teams maintain quality.

The principles in this post apply across programming languages. The examples use C# because it is the language I’m most familiar with.

AI-assisted coding has changed the cost of developing software. For many tasks, that is genuinely useful. But review and maintenance still take time, and this creates a mismatch worth paying attention to.

Consider a hypothetical pull request for a payment retry feature. The assistant produces a large, polished change quickly. The code compiles, and the tests pass, but the reviewer has limited time to understand every retry path and helper method.

Weeks later after this feature get's released to production a customer accidently get's charged twice. How could I have missed this?

The risk is not that AI always writes bad code. It is that producing code has become cheaper while understanding and verifying it still require engineering time. When the volume of changes grows beyond the team’s review capacity, defects become easier to miss.

AI coding tools can help developers finish certain tasks faster. In one controlled GitHub experiment, professional developers using Copilot completed a specific programming task 55% faster on average. That result applies to one task, and does not guarantee that every feature or the whole development lifecycle will be 55% faster.<sup>1</sup>

The risk appears when code production speeds up but review capacity does not.

Illustrative, not measured data

Code arriving:  ███████████████
Careful review: ██████
                └── The rest waits or gets less attention

When a pull request is too large to understand in the time available, reviewers can end up checking whether it looks right instead of working through what it actually does. Polished code can make this harder: familiar patterns and tidy names create confidence, but they don’t prove that the design is correct or that edge cases are handled.

DORA’s 2024 research describes a similar tension at the delivery level. AI adoption was associated with reported gains in individual productivity, flow, and job satisfaction, alongside negative effects on software delivery stability and throughput.<sup>2</sup>

That is not proof that AI inevitably lowers code quality. It is a reason to measure what happens after code is generated, not just how quickly it appeared.

Imagine asking an assistant:

Retry a failed payment request up to three times.

It might generate something like this:

public async Task ChargeOrderAsync(Order order)
{
    for (var attempt = 0; attempt < 3; attempt++)
    {
        try
        {
            await _paymentGateway.ChargeAsync(order.Amount);
            return;
        }
        catch (Exception) when (attempt < 2)
        {
            // Retry after a failure.
        }
    }
}

The loop is easy to follow. It may even pass tests for a successful charge and a request that fails immediately.

But what if the payment gateway processes the charge, then the response times out before the application receives it? The next attempt may charge the customer again.

The problem is not that the code is messy. A key behavior, what happens when the outcome is unknown, was never made explicit.

A safer design usually needs a stable idempotency key for the logical order, reused across retries. The exact API depends on the payment provider, but the idea might look like this:

public async Task ChargeOrderAsync(Order order)
{
    var idempotencyKey = $"order:{order.Id}";

    for (var attempt = 0; attempt < 3; attempt++)
    {
        try
        {
            await _paymentGateway.ChargeAsync(order.Amount, idempotencyKey);
            return;
        }
        catch (Exception exception)
            when (IsRetryable(exception) && attempt < 2)
        {
            continue;
        }
    }
}

This is illustrative pseudocode, not a drop-in implementation. A real change still needs to follow the provider’s idempotency rules and define which errors are safe to retry.

Google’s code review guidance recommends looking beyond whether code “works”: reviewers should consider design, complexity, tests, context, and whether they understand every line.<sup>3</sup> Those questions matter especially when a patch was generated quickly and looks convincing at a glance.

Start with questions, not implementation:

Every AI assistant usually has a plan mode. Ask the AI to write out a plan for the feature you want to implement and make sure questions like these are answered before any code gets written.

Split work into focused changes: one for the behavior, another for an unrelated cleanup, and another for follow-up documentation.

A small pull request gives a reviewer a better chance to understand how the pieces fit together. It also makes it easier to identify what caused a failure later. DORA recommends reducing batch size for similar reasons: smaller changes are easier to reason about and recover from.<sup>4</sup>

For the payment example, useful tests might cover:

Don’t stop at “the tests pass.” Ask whether a test would fail if the bug you’re worried about came back. Tests are code too, and a test that repeats the implementation’s assumptions may confirm the wrong behavior.

The developer submitting a pull request should be able to explain what every changed part does, why it belongs, and what risks remain.

AI can help summarize a diff or suggest review questions, but that summary should not replace reading the diff. Treat it as a map, not proof that you have visited every place on it.

Lines generated and tasks started are easy to count; they don’t tell you whether the change helped users or became expensive to maintain.

Look at several signals together: review time, rework, changes that need to be rolled back or urgently fixed, and whether delivery is becoming more or less stable. DORA’s guidance cautions against relying on one metric or turning measurements into targets that teams feel pressured to game.<sup>4</sup>

These are delivery signals, not direct measures of code quality. Pair them with code review and testing practices to understand what is actually improving or getting worse.

Define behavior
      ↓
Ask for a plan
      ↓
Generate one small change
      ↓
Test edge cases and inspect the diff
      ↓
Human review and ownership
      ↓
Merge, monitor, and learn

Before approving, ask:

If the honest answer to “Do I understand this?” is no, ask for clarification or a smaller change before approving it.

AI can shorten the distance between an idea and a working first draft. It can also make it easier for a team to accumulate changes that nobody has had time to understand.

The useful goal is to keep the size and arrival rate of changes within the team’s ability to review them carefully. If your team is trying AI tools, start with one small change and look at more than how quickly it ships. Check the tests, review effort, and cost of maintaining the result too.

GitHub’s experiment recruited 95 professional developers to complete a JavaScript HTTP server task. The result is specific to that experiment and doesn’t establish that all development work gets faster. ↩

DORA reports associations from its research; those findings don’t establish that AI alone caused changes in delivery performance. ↩

Google’s guidance was written for code review generally, not specifically for AI-generated changes. ↩

DORA presents these as delivery performance signals, not direct measures of code quality. ↩

── more in #ai-tools 4 stories · sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-ai-can-lead-to-a…] indexed:0 read:6min 2026-09-30 · —