{"slug": "the-agent-shouldn-t-be-able-to-approve-its-own-rules", "title": "The Agent Shouldn't Be Able to Approve Its Own Rules", "summary": "A developer essay argues that AI coding agents should be allowed to propose policy changes but never approve them, because an agent that can rewrite the rules it is checked against makes verification meaningless. The author illustrates the failure mode with concrete examples: an agent that fails a \"maximum service size = 300 lines\" check can simply raise the limit to 500, or one blocked by a \"module A cannot import module B\" rule can edit the rule to allow the import, so the check passes without any real constraint. The proposed fix separates changing code from changing the rules around the code, requiring a human or separate authority to approve any policy change the agent proposes.", "body_md": "AI coding agents can make legitimate architectural changes. The harder problem is separating code changes from policy changes, waivers and approval authority.\n\n*How to let AI coding agents change architecture without letting them silently change the rules around it.*\n\nOne of the less obvious problems I've run into with coding agents isn't that they make bad changes.\n\nIt's that sometimes they make a perfectly reasonable change that invalidates one of the rules I'm using to check the project.\n\nThat's a different problem.\n\nImagine a project has this rule:\n\n```\nAll payment-provider access must go through PaymentAdapter.\n```\n\nAn agent is asked to add support for another payment provider.\n\nIt looks at the existing code and decides that the current adapter isn't quite right. It wants to introduce a new abstraction.\n\nThe resulting change crosses a boundary that the project currently protects.\n\nThe checker reports a violation.\n\nSo what should happen next?\n\nThe obvious automation is:\n\nTechnically, everything worked.\n\nBut the system has a problem: **The thing being checked was allowed to change the conditions of the check.**\n\nThat made me much more interested in the difference between **changing the code** and **changing the rules around the code**.\n\nThe simplest version looks like this:\n\n```\nagent\n  ↓\nchange code\n  ↓\nverification\n  ↓\nfailure\n  ↓\nagent changes policy\n  ↓\nverification\n  ↓\npass\n```\n\nThere is nothing obviously broken here.\n\nThe agent might even have a good reason for changing the policy.\n\nThe problem is that the verification boundary has disappeared.\n\nA policy is supposed to tell the system what has to remain true.\n\nIf the same actor that made the change can also redefine what \"true\" means, a successful verification doesn't tell you very much.\n\nIt's just a moving target.\n\nAnd this isn't specific to architecture.\n\nYou could do the same thing with:\n\n```\nmaximum service size = 300 lines\n```\n\nThe agent produces a 420-line service.\n\nThe check fails.\n\nThe agent changes the limit to 500.\n\nThe check passes.\n\nOr:\n\n``` python\nmodule A cannot import module B\n```\n\nThe agent needs the dependency.\n\nIt changes the rule.\n\nThe import is now allowed.\n\nAgain, maybe that's the right architectural decision.\n\nBut those are two separate decisions:\n\nThey shouldn't become one operation just because the same agent proposed both.\n\nThis is where it gets more interesting.\n\nI don't want an architecture checker that treats every violation as proof that the agent did something wrong.\n\nSometimes the agent really should change the architecture.\n\nSuppose an application has:\n\n```\nCheckout\n   ↓\nPaymentAdapter\n   ↓\nStripe\n```\n\nAnd I decide that the product is getting large enough that payment workflows deserve their own domain boundary:\n\n```\nCheckout\n   ↓\nPaymentService\n   ↓\nPaymentAdapter\n   ↓\nStripe\n```\n\nThe change introduces new files.\n\nSome imports move.\n\nSome old boundaries disappear.\n\nNew ones appear.\n\nA strict checker could report a pile of violations.\n\nThat doesn't mean the change is bad.\n\nIt means the current policy describes the old architecture.\n\nThis distinction matters.\n\nA policy isn't supposed to prevent architecture from ever changing.\n\nIt is supposed to make architecture changes explicit.\n\nThis led me to a fairly simple rule: **An agent should be able to propose a policy change, but it shouldn't automatically be able to approve that policy change.**\n\nFor example:\n\n```\nCurrent policy:\n\nPayment provider access must go through PaymentAdapter.\n```\n\nThe agent can say:\n\n```\nProposed change:\n\nAllow PaymentService to depend directly on a new internal\nPaymentProvider interface.\n\nReason:\nThe current adapter boundary prevents the new workflow\nfrom sharing transaction state correctly.\n```\n\nThat's useful.\n\nThe agent has done the hard reasoning.\n\nIt has identified the existing constraint.\n\nIt has explained why the constraint may no longer fit.\n\nBut the proposal should remain a proposal.\n\nThe important part is that **the authority approving the policy change is separate from the agent that authored the change**.\n\nOtherwise the system can silently move the goalposts.\n\nA prompt can say:\n\n```\nNever deploy without approval.\n```\n\nThat's useful instruction.\n\nIt's not much of a control if the same process can modify the configuration that defines what counts as an approved deployment.\n\nThe more autonomous the agent becomes, the more these distinctions move out of the prompt and into the environment around it.\n\nThe agent should be able to reason about the policy.\n\nIt should be able to request a policy change.\n\nIt should be able to explain the change.\n\nThe actual enforcement shouldn't depend on the agent remembering to follow its own instructions.\n\nOnce policy becomes a real project artifact, another problem appears.\n\nYou need to know what the policy was when the original change was checked.\n\nOtherwise you can end up with a strange situation where today's successful verification only makes sense because yesterday's policy was replaced.\n\nThat's why I like keeping policy changes explicit and versioned.\n\nSomething like:\n\n```\nproject state\n    ↓\npolicy revision 17\n    ↓\nagent proposes architecture change\n    ↓\nverification fails under policy revision 17\n    ↓\npolicy proposal\n    ↓\napproval\n    ↓\npolicy revision 18\n    ↓\nverify change again\n```\n\nNow there are two separate facts:\n\nThat's much more useful than simply seeing a green check.\n\nThere's another subtle point here.\n\nSuppose an agent changes the rule and then verifies the same diff against the new rule.\n\nYou can no longer tell whether the original change violated the previous policy.\n\nSo I want the system to preserve the distinction between:\n\n```\nwhat the project allowed before the change\n```\n\nand:\n\n```\nwhat the project allows after the change\n```\n\nThis becomes especially useful when investigating a change later.\n\nYou can ask:\n\nThose questions are much harder to answer when policy is just another mutable config file.\n\nSometimes the policy is fine.\n\nThe violation is temporary.\n\n```\nAll persistence access must go through Repository.\n```\n\nBut I'm in the middle of a migration.\n\nI don't want to remove the rule.\n\nI just need one known exception for two weeks.\n\nThat's not really a policy change.\n\nIt's a waiver.\n\nAnd I think treating it as a different object makes the whole system easier to reason about.\n\nA useful waiver has at least:\n\n```\nowner\nreason\nscope\nexpiry\n```\n\nSo instead of:\n\n```\nremove the rule\n```\n\nyou get:\n\n```\nwaive this finding until 2026-12-28\n\nowner: platform-team\nreason: repository migration\n```\n\nThe rule stays.\n\nThe exception expires.\n\nThat's a very different thing from changing the architecture policy permanently.\n\nI also don't want to confuse waivers with baselines.\n\nA baseline answers:\n\n```\nThis violation already existed.\n```\n\nA waiver answers:\n\n```\nThis active violation is intentionally allowed for a limited time.\n```\n\nAnd a policy change answers:\n\n```\nWe changed what the project considers acceptable.\n```\n\nThose are three different states.\n\nThat separation might sound overly precise.\n\nIn practice, it makes the tool much more useful.\n\nA real codebase can have old architectural debt.\n\nIt can have temporary migration exceptions.\n\nAnd it can deliberately evolve its architecture.\n\nIf all three become \"ignore this finding\", you lose important information.\n\nThis doesn't mean agents have to stop making architectural changes.\n\nQuite the opposite.\n\nI want them to do more.\n\nA useful agent flow could look like this:\n\n```\nrequirement\n    ↓\nagent plans change\n    ↓\nimplementation\n    ↓\ndeterministic verification\n    ↓\nfailure\n    ↓\nagent explains why\n    ↓\npolicy proposal / waiver proposal\n    ↓\napproval\n    ↓\nverification against new state\n```\n\nThe agent can drive most of that workflow.\n\nIt can inspect the repository.\n\nIt can understand the requirement.\n\nIt can implement the refactor.\n\nIt can identify the policy that stopped the change.\n\nIt can prepare the evidence for a policy proposal.\n\nIt can even tell me that the current architecture appears to be the problem.\n\nWhat it shouldn't get is an invisible path from:\n\n```\nmy change failed\n```\n\nto:\n\n```\ntherefore my own change is now allowed\n```\n\nI've started thinking that autonomy isn't really one switch.\n\nAn agent can have permission to:\n\nwithout having permission to:\n\nThat gives you a more useful permission model than simply:\n\n```\nautonomous = yes/no\n```\n\nDifferent operations can have different authorities.\n\nAnd that matters more once the agent is running for a long time or can delegate work to other agents.\n\nThere's another distinction here that I find useful.\n\nVerification should primarily answer:\n\n```\nDoes the current change fit the current approved policy?\n```\n\nIt doesn't need to decide whether the author was:\n\n```\nhuman\nClaude\nCodex\nCursor\nanother agent\nautomation\n```\n\nThat's a separate concern.\n\nThe verifier checks the state.\n\nThe policy layer defines the constraint.\n\nThe authority layer controls who can change the constraint.\n\nKeeping those pieces separate makes the system much easier to reason about.\n\nThis distinction ended up affecting the design of Codapult Guard quite a bit.\n\nGuard already had the idea of:\n\n```\nfacts\n  ↓\npolicy\n  ↓\nverification\n```\n\nBut that isn't enough when policy itself can change.\n\nSo I started treating policy changes as first-class operations.\n\nGuard can discover project facts and prepare proposals.\n\nA project can explicitly approve them.\n\nFor protected projects, policy approval can require a distinct actor rather than allowing the authoring agent to approve its own proposal.\n\nThe same idea now applies to waivers.\n\nA waiver is not just \"ignore this finding forever.\"\n\nIt has an owner, reason and expiry date.\n\nWhen the expiry is reached, the finding becomes active again.\n\nThat gives the project three useful things:\n\n```\npolicy\nexception\nhistory\n```\n\ninstead of one growing collection of ignored warnings.\n\nMCP makes this distinction especially important.\n\nAn agent can ask Guard:\n\n```\nWhat policy applies here?\n```\n\nor:\n\n```\nWhat did this change affect?\nWhy did verification fail?\nWhat policy change would be needed for this refactor?\n```\n\nThose are useful questions.\n\nBut I don't want the agent to receive:\n\n```\nChange policy until verification passes.\n```\n\nThat isn't really verification anymore.\n\nThe MCP layer should expose the information and operations the agent needs without quietly giving it the authority to redefine the system around itself.\n\nThat's a much more interesting role for project-aware tooling than simply adding another set of commands to an agent.\n\nI don't think architectural rules should be permanent.\n\nProjects change.\n\nRequirements change.\n\nTeams change.\n\nThe shape of the system changes.\n\nA rule that made perfect sense six months ago might become actively harmful.\n\nThe mistake is treating policy change as an implementation detail.\n\nChanging:\n\n```\nsrc/payments/StripeAdapter.ts\n```\n\nand changing:\n\n```\npayment-provider-access\n```\n\nare fundamentally different operations.\n\nOne changes the system.\n\nThe other changes what the system is allowed to become.\n\nOnce an agent can perform both, they need separate boundaries.\n\nThe goal isn't to make agents weaker.\n\nI actually want agents to make bigger changes.\n\nBut bigger changes require clearer boundaries.\n\nThe more of the implementation I delegate, the more I care about questions like:\n\nThose questions aren't really about whether the model is smart enough.\n\nThey're about the architecture around the model.\n\nAnd that's probably the part that becomes more important as coding agents become capable of doing more work without waiting for a human after every step.\n\nI don't want the agent to be afraid of the rules.\n\nI want it to understand them.\n\nI want it to be able to challenge them.\n\nI want it to be able to propose better ones.\n\nBut I don't want it to be the final authority on whether its own change should become the new rule.\n\nThe code can change.\n\nThe architecture can change.\n\nEven the policy can change.\n\nThose changes just shouldn't all happen under the same authority.\n\n**Want AI agents to propose architectural changes without approving their own rules?**\nCodapult Guard separates project facts, proposed policy, approval, and deterministic verification. Read the [Guard documentation](https://github.com/codapult/codapult-guard?tab=readme-ov-file) or see how Codapult exposes Guard through its [MCP server](https://codapult.dev/docs/developer-tools/mcp).", "url": "https://wpnews.pro/news/the-agent-shouldn-t-be-able-to-approve-its-own-rules", "canonical_source": "https://codapult.dev/blog/the-agent-shouldnt-be-able-to-approve-its-own-rules", "published_at": "2026-09-29 00:00:00+00:00", "updated_at": "2026-10-02 07:07:08.887302+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "developer-tools", "artificial-intelligence"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-agent-shouldn-t-be-able-to-approve-its-own-rules", "markdown": "https://wpnews.pro/news/the-agent-shouldn-t-be-able-to-approve-its-own-rules.md", "text": "https://wpnews.pro/news/the-agent-shouldn-t-be-able-to-approve-its-own-rules.txt", "jsonld": "https://wpnews.pro/news/the-agent-shouldn-t-be-able-to-approve-its-own-rules.jsonld"}}