{"slug": "loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents", "title": "Loop Engineering in Practice: Six Feedback Loops for AI Coding Agents", "summary": "An engineering team describes six feedback loops spanning planning, coding, testing, review, release, and incidents that turn an AI coding agent's corrections into durable checks across the delivery pipeline. The approach emerged from an account-deletion feature where the agent's passing tests missed active subscriptions and shared-project ownership, prompting the team to encode product decisions as acceptance tests and require executable environments and end-to-end verification. The team also defined stop conditions so that repeated identical failures return evidence to a human engineer rather than triggering further edits.", "body_md": "This is loop engineering in practice: six feedback loops that turn an AI coding agent's corrections into lasting checks across the delivery pipeline.\n\nWe were recently implementing a feature that allowed users to delete their accounts. A coding agent had implemented the endpoint, added tests, and produced a pull request. The tests passed, but review showed that active subscriptions remained untouched. Further testing revealed that deleting the owner also broke access to shared objects.\n\nEach discovery gave us information about what the implementation had missed. We wanted to create a proper process to utilise this information: a process to correct the immediate change, then preserve the expectation so it influenced subsequent work.\n\nThe work brought us back to six feedback loops across planning, coding, testing, review, release, and incidents. Designing the cycles in which an agent acts, observes the result, and revises its approach now has a name, [loop engineering](https://addyosmani.com/blog/loop-engineering/). The loops below extend that idea beyond the coding agent to the whole delivery pipeline, and this is where I feel most engineering leaders should be focusing when designing an AI-first engineering pipeline.\n\nOur requirement, “Allow users to delete their accounts”, had described an intention but left several product decisions open.\n\nShould deletion happen immediately? What would happen to an active subscription? Could the sole owner of a shared project delete their account?\n\nWe agreed that subscriptions had to be cancelled and ownership of shared projects transferred before deletion. Those decisions gave us specific expectations to test:\n\nAI helped us inspect the existing implementation and surface unanswered questions and edge cases. Resolving the intended behaviour still required product context and engineering judgement.\n\nWe closed the planning loop by baking those answers into the specification and then converting them to acceptance tests. Leaving them in a chat would have made them easy for the next implementation to miss.\n\nDuring implementation, the agent needed a reliable way to run the application and inspect failures. We made build output, relevant tests, and runtime context available so it could examine where its changes broke.\n\nIn one correction, the deletion endpoint had called a service with the wrong argument type. We gave the agent the actual error. It inspected the service contract, corrected the call, and reran the failing check.\n\nThe rerun mattered because the edit alone did not establish that the problem had been resolved.\n\nAnthropic’s work on [long-running coding agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents) describes failures where agents declared features complete without adequate testing. In its experiments, providing an executable environment and prompting end-to-end verification helped expose problems that code inspection had missed.\n\nWe also defined when to stop iterating. If repeated attempts produced the same failure, the agent needed to return the evidence and its attempted fixes for review. Any additional edit needs a proper reason vetted by an engineer.\n\nOur initial tests had covered the scenarios we had thought to encode. Further testing exposed another gap: a subscription lookup failure due to a malformed lookup id was being treated as “no active subscription”. The application consequently allowed deletion when it could not establish eligibility.\n\nWe agreed that deletion should be blocked in that situation, then added a regression test to reproduce the failure. We checked that it failed against the broken implementation and passed after the correction.\n\nWe also reviewed the assertions. Just asserting that the endpoint returned an error would have been incomplete if it had already deleted the account. We needed to inspect the resulting state. That assertion gap deserves its own treatment: [AI in Software Testing: Why Generated Tests Miss Bugs](https://nulltensor.com/posts/ai-generated-test-coverage/?utm_source=devto&utm_medium=crosspost&utm_campaign=ai-generated-test-coverage).\n\nAI helped construct the reproduction and implement the fix. We need to judge whether the test captures the failure accurately and whether its expected outcome matches the business rule.\n\nReview surfaced a permission issue that our existing checks had missed. An ownership check trusted a user ID supplied in the request.\n\nFixing the endpoint addressed the immediate issue. We also considered where the same mistake could recur and how to make the correction available to subsequent work.\n\nWe added a test for cross-account access, used an established authorisation helper, and documented when to use it in the repository guidance.\n\nThese serve different purposes. The guidance helps the agent choose an approach while the test checks a specific outcome.\n\nWe did not turn every review comment into a permanent rule. We focused on corrections that expressed recurring constraints and made them available before the agent’s next implementation. A critical comment buried in a merged pull request becomes a weak dependency for future work until it is formalised.\n\nBefore release, we defined the evidence we needed to decide whether rollout should continue.\n\nGoogle’s guidance on [canary releases](https://sre.google/workbook/canarying-releases/) explains how exposing a change to limited traffic and comparing its behaviour with a control set can inform that decision.\n\nFor account deletion, request success alone was insufficient. We also needed visibility into background cleanup and unexpected failures in eligibility checks.\n\nWe defined pause conditions, assigned different engineers responsibility across them, and mandated recovery before further deployment. Rolling back application code would not restore data already deleted.\n\nAI constantly helped us examine logs and metrics, but the rollout still needed explicit decision rules vetted by engineers. Scaling the canary becomes a direct function of the observed behaviour, where the function rules are defined by us.\n\nA later incident exposed a sequence we had missed: an account had become eligible for deletion, a subscription had been created concurrently, and deletion had proceeded using the earlier eligibility result.\n\nWe reconstructed the sequence and examined the contributing conditions. AI helped organise the evidence and propose explanations, which we checked against the observed behaviour.\n\nThe follow-up included a concurrency test and a change to how eligibility and deletion were coordinated. Improving detection helps but it addresses another part of the problem because an alert alone does not have prevent the sequence.\n\nGoogle’s [postmortem guidance](https://sre.google/sre-book/postmortem-culture/) emphasises understanding contributing causes and implementing preventive actions. We assigned owners to the follow-up work and defined the evidence needed to verify that it addressed the failure.\n\nThat evidence collected from these production discoveries should be properly routed back to the planning, implementation, and testing phases. It changes the expectation from each phase of the next release. AI helped us brainstorm which phase each evidence can roughly be attributed to, but the final decision of ownership lies with the engineering leader running the whole cycle.\n\nAcross these six stages, we needed to make four things explicit: the feedback signal, who received it, what action it could trigger, and how we would verify the correction.\n\nAn agent could act on feedback within a run. Subsequent runs needed the relevant tests, instructions, and evidence made available again. We had to design that continuity into the workflow.\n\nThe most low hanging fruit to implement from this entire process is to start with a correction reviewers keep making or a failure that has escaped more than once. Trace where it becomes visible today, then build a check that brings it forward. Use the next relevant change to see whether the loop actually works.\n\n*Originally published on [nulltensor.com](https://nulltensor.com/posts/six-feedback-loops-ai-first-engineering-pipeline/?utm_source=devto&utm_medium=crosspost&utm_campaign=six-feedback-loops-ai-first-engineering-pipeline).*", "url": "https://wpnews.pro/news/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents", "canonical_source": "https://dev.to/rss_holmes/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents-1a2n", "published_at": "2026-10-01 03:00:00+00:00", "updated_at": "2026-10-01 03:16:39.900480+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops", "artificial-intelligence"], "entities": ["Anthropic", "Addy Osmani"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents", "markdown": "https://wpnews.pro/news/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents.md", "text": "https://wpnews.pro/news/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents.txt", "jsonld": "https://wpnews.pro/news/loop-engineering-in-practice-six-feedback-loops-for-ai-coding-agents.jsonld"}}