{"slug": "how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget", "title": "How I Stopped My AI Coding Agent's Scope Creep With a Diff Budget", "summary": "A developer running a fully autonomous coding-agent system curbed agent scope creep by introducing a per-task \"diff budget\": each task spec declares allowed path globs plus max_files and max_changed_lines limits, a Claude Code PreToolUse hook blocks out-of-scope edits before they land and feeds the reason back to the model, and a final diff check enforces the size cap. Drive-by improvements are diverted into a \"parking lot\" that the orchestrator turns into follow-up tasks, with the developer arguing that instructions describe intent while only a mechanism provides a guarantee.", "body_md": "My autonomous coding agent kept turning small tasks into sprawling diffs: ask for a bug fix, get a bug fix plus a renamed helper, a reformatted file, and a \"quick cleanup\" nobody asked for. I fixed it with a **diff budget**: every task declares which paths it may touch and how big the change is allowed to be, a hook blocks out-of-scope edits before they happen, and a \"parking lot\" turns the agent's drive-by ideas into follow-up tasks instead of surprise changes. Here's the design, the code, and 5 lessons from running it.\n\nI run a fully autonomous implementation system: an orchestrator module hands tasks to parallel implementation agents, and they work through a backlog while I'm doing something else (often sleeping 😴).\n\nThe agents are good at the tasks. That was never the problem. The problem was everything they did *in addition to* the tasks.\n\nA typical example: the task says \"fix the off-by-one in the pagination helper.\" The agent fixes it. Then, since it's already in the file, it:\n\nEvery one of those changes is defensible on its own. Together they turn a five-line fix into a diff that touches a dozen files, and that hurts in three specific ways:\n\nMy first attempt was the obvious one: add a line to the instructions. \"Only change what the task requires. Do not refactor unrelated code.\"\n\nIt helped a little, and not nearly enough. The agent doesn't experience its cleanup as \"unrelated.\" From inside the task, renaming a confusing variable feels like part of doing a good job. I was asking a model to resist something it considers a virtue, using a sentence. That's a losing fight.\n\nInstructions describe intent. If you need a guarantee, you need a mechanism.\n\nSo I stopped trying to persuade the agent and built a fence instead.\n\nThe design has three pieces: a budget declared per task, a hook that enforces the path part *before* an edit lands, and a check on the final diff that enforces the size part. Plus one escape valve, which turned out to be the piece that made everything else work.\n\n``` php\nflowchart TD\n    A[Orchestrator creates task] --> B[Task spec with scope + budget]\n    B --> C[Implementation agent starts]\n    C --> D{Edit inside allowed paths?}\n    D -- yes --> E[Edit applied]\n    D -- no --> F[Blocked: agent told to use parking lot]\n    F --> G[Parking lot note written]\n    E --> H{Final diff within budget?}\n    H -- yes --> I[Commit + hand to review]\n    H -- no --> J[Returned to agent: trim or request extension]\n    G --> K[Orchestrator files follow-up tasks]\n```\n\nWhen the orchestrator module creates a task, the spec now carries a `scope` block. It's deliberately boring:\n\n```\n{\n  \"id\": \"task-0412\",\n  \"goal\": \"Fix off-by-one in pagination helper\",\n  \"scope\": {\n    \"allowed_paths\": [\"src/pagination/**\", \"tests/pagination/**\"],\n    \"max_files\": 4,\n    \"max_changed_lines\": 80\n  }\n}\n```\n\nThree numbers and a list of globs. The planner step sets them when it writes the task, based on the task type: bug fixes get a tight budget, feature work gets a wider one, and refactors are their own task type with an explicitly large budget. That last part matters. I'm not against refactoring. I'm against refactoring *smuggled inside* something else.\n\nClaude Code (I'm on v2.1.x) supports hooks that run before a tool call. A `PreToolUse` hook receives the tool call as JSON on stdin, and if it exits with code 2, the call is blocked and whatever the hook wrote to stderr is fed back to the model.\n\nThat feedback channel is the whole trick. The agent doesn't just hit a wall. It gets told why, and what to do instead.\n\n``` bash\n#!/usr/bin/env python3\n\"\"\"PreToolUse hook: block edits outside the current task's allowed paths.\"\"\"\nimport fnmatch\nimport json\nimport os\nimport sys\n\ncall = json.load(sys.stdin)\nif call.get(\"tool_name\") not in (\"Edit\", \"Write\"):\n    sys.exit(0)\n\nwith open(os.environ[\"TASK_SPEC\"]) as f:\n    scope = json.load(f)[\"scope\"]\n\nroot = call.get(\"cwd\", os.getcwd())\ntarget = os.path.relpath(call[\"tool_input\"][\"file_path\"], root)\n\nif any(fnmatch.fnmatch(target, pattern) for pattern in scope[\"allowed_paths\"]):\n    sys.exit(0)\n\nprint(\n    f\"BLOCKED: {target} is outside this task's scope \"\n    f\"({', '.join(scope['allowed_paths'])}).\\n\"\n    \"If this change is required to finish the task, stop and request a \"\n    \"scope extension with a one-line reason.\\n\"\n    \"If it is an improvement you noticed, append it to the parking lot \"\n    \"file and keep going.\",\n    file=sys.stderr,\n)\nsys.exit(2)\n```\n\nRegistering it is a few lines in the settings file:\n\n```\n{\n  \"hooks\": {\n    \"PreToolUse\": [\n      {\n        \"matcher\": \"Edit|Write\",\n        \"hooks\": [{ \"type\": \"command\", \"command\": \"python3 hooks/scope_guard.py\" }]\n      }\n    ]\n  }\n}\n```\n\nOne honest caveat: this hook only covers the file-editing tools. An agent with shell access can still modify files through a shell command, so the hook is a guardrail, not a security boundary. That's why the next check exists.\n\nPath scoping stops the agent from wandering. It doesn't stop it from rewriting everything *inside* the allowed directory. So before a task's work is committed, a small script measures the actual diff against the budget:\n\n``` bash\n#!/usr/bin/env python3\n\"\"\"Compare the working tree diff against the task's budget.\"\"\"\nimport fnmatch\nimport json\nimport subprocess\nimport sys\n\nscope = json.load(open(sys.argv[1]))[\"scope\"]\n\nout = subprocess.run(\n    [\"git\", \"diff\", \"--numstat\", \"HEAD\"],\n    capture_output=True, text=True, check=True,\n).stdout\n\nfiles, lines, outside = 0, 0, []\nfor row in out.strip().splitlines():\n    added, deleted, path = row.split(\"\\t\")\n    files += 1\n    if added != \"-\":  # binary files report \"-\"\n        lines += int(added) + int(deleted)\n    if not any(fnmatch.fnmatch(path, p) for p in scope[\"allowed_paths\"]):\n        outside.append(path)\n\nproblems = []\nif outside:\n    problems.append(f\"out-of-scope files: {', '.join(outside)}\")\nif files > scope[\"max_files\"]:\n    problems.append(f\"{files} files changed (budget {scope['max_files']})\")\nif lines > scope[\"max_changed_lines\"]:\n    problems.append(f\"{lines} lines changed (budget {scope['max_changed_lines']})\")\n\nif problems:\n    print(\"OVER BUDGET: \" + \"; \".join(problems))\n    sys.exit(1)\nprint(f\"within budget: {files} files, {lines} lines\")\n```\n\nBecause this reads the real diff from git, it catches everything regardless of how the file got modified, including the shell-command route the hook can't see.\n\nWhen the check fails, the task isn't thrown away. The output goes back to the agent with a simple choice: trim the diff down to what the task needs, or request an extension.\n\nMy first version only had the fence, and it had a nasty side effect: the agent would notice a real problem outside its scope, get blocked, and then the observation simply vanished. Some of those observations were valuable. A genuinely broken neighbor module is something I want to know about.\n\nSo each task gets a parking lot: a plain Markdown file where the agent appends things it noticed but wasn't allowed to touch.\n\n```\n## Parking lot: task-0412\n\n- src/search/cursor.py has the same off-by-one pattern as the pagination\n  helper. Likely the same bug. Suggested task: bug fix, tight budget.\n- Variable names in src/pagination/window.py are misleading (`end` is\n  inclusive). Suggested task: refactor, low priority.\n```\n\nWhen the task finishes, the orchestrator reads the parking lot and files each entry as a candidate task in the backlog, with its own scope and budget. The cleanup still happens. It just happens as its own small, reviewable, revertable change.\n\nThis changed the agent's behavior more than the fence did. Once it had somewhere legitimate to put its ideas, it stopped trying to sneak them in.\n\nSometimes the budget is simply wrong. The fix really does need a change to a shared type definition two directories over. For that, the agent can request an extension: it stops, writes one line explaining which path it needs and why, and the orchestrator decides.\n\nSmall extensions that stay inside the same package get approved automatically. Anything touching shared or high-traffic areas waits for me. In practice most requests are reasonable, and the ones that aren't are usually a sign the task was specified badly in the first place, which is useful to learn.\n\n**1. Scope creep is a virtue in the wrong place.** The agent isn't misbehaving when it tidies up. It's doing what a conscientious engineer does. Treating it as disobedience leads you to write angrier instructions. Treating it as misrouted effort leads you to build a parking lot.\n\n**2. Block with an explanation, never silently.** A hook that just says \"denied\" makes the agent try the same thing a different way. A hook that says \"this is out of scope, here are your two options\" gets a sensible next move almost every time. The error message is a prompt. Write it like one.\n\n**3. Give the agent a legitimate outlet.** Every constraint needs an answer to \"okay, then what do I do with this?\" Without the parking lot, I was throwing away real findings. With it, the constraint became a routing rule instead of a muzzle.\n\n**4. Budgets should be set by task type, not by gut.** A single global limit is always wrong for somebody: too loose for bug fixes, too tight for features. Three or four task types with different defaults covered nearly everything I needed.\n\n**5. Small diffs are a feature of the system, not a nicety.** Review speed, parallel agents not colliding, clean reverts: all of these depend on changes being small and single-purpose. If you run agents unattended, diff size is the thing to protect first.\n\nIf your coding agent keeps handing you diffs three times bigger than the task, don't write a sterner instruction. Declare a budget, enforce it with a hook, check the real diff, and give the agent a parking lot for its good ideas. You can wire up a first version with the two scripts above in an afternoon.\n\nIf this was useful, **follow me here on Dev.to**. I write regularly about building and running a fully autonomous implementation system with Claude Code. And I'd love to hear in the comments: how do you keep your agent's diffs small? 🚀", "url": "https://wpnews.pro/news/how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget", "canonical_source": "https://dev.to/yureki_lab/how-i-stopped-my-ai-coding-agents-scope-creep-with-a-diff-budget-47lj", "published_at": "2026-10-03 14:31:31+00:00", "updated_at": "2026-10-03 14:37:54.864580+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "mlops"], "entities": ["Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget", "markdown": "https://wpnews.pro/news/how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget.md", "text": "https://wpnews.pro/news/how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget.txt", "jsonld": "https://wpnews.pro/news/how-i-stopped-my-ai-coding-agent-s-scope-creep-with-a-diff-budget.jsonld"}}