Your coding agent is following rules that are no longer true. Here is the fix. A developer argues that stale agent instruction files like AGENTS.md are more dangerous than missing rules because they make coding agents confidently follow outdated commands, and proposes a CI doc-drift check that compares a doc's last commit time against the paths it covers, plus a discard-and-retry procedure to verify a new rule actually changes behavior. The writeup notes the check requires a full git checkout (fetch-depth: 0) or every doc will appear freshly updated. The most dangerous agent bug is not a model that invents an API. It is an instruction file that is still true enough to be trusted, and stale enough to be wrong. Your AGENTS.md says the tests run with npm test . Last month a teammate moved them behind a build step. The agent reads the rule, trusts it, runs the old command, and reports a clean pass that never actually ran. No error is thrown. The stale rule is, by design, silent. That is worse than having no rules at all. No rules make the agent conservative. A stale rule makes it confident. An instruction file goes stale without failing. It references a path that moved, a command that changed, or a convention the team abandoned. Nothing checks whether the rule is still true, because the rule is plain prose with no owner and no schedule. The result is a trust asymmetry: the file looks authoritative, so the agent follows it faithfully and confidently does the wrong thing. The maintainer has no signal that the guidance drifted, because the file does not throw on mismatch the way code does. The fix starts with a change of category. AGENTS.md is not documentation you write once and forget. It is a living input to a system that consumes it on every run. Treat it on the same maintenance cadence as your README or your API docs — which is to say, it needs a review trigger. Two mechanisms keep it honest. One catches drift after the fact. The other verifies that a rule actually changes behavior before you trust it. Give each agent-visible file a covers: line that lists the paths it describes. Then add a CI job that compares the last commit time of each doc against the last commit time of the paths it covers. scripts/check doc drift.sh fails when a doc is older than its subject by more than MAX LAG DAYS MAX LAG DAYS=14 for doc in docs/agents/ .md; do covers=$ grep -oP '^covers:\s \K.+' "$doc" || true -z "$covers" && continue doc ts=$ git log -1 --format=%ct -- "$doc" for target in $covers; do target ts=$ git log -1 --format=%ct -- "$target" lag=$ target ts - doc ts / 86400 if "$lag" -gt "$MAX LAG DAYS" ; then echo "STALE: $doc is $lag days behind $target"; exit 1 fi done done echo "Docs OK" The one gotcha that will bite you: the git checkout in CI must use fetch-depth: 0 . Otherwise git log sees a single commit since the shallow clone, every doc looks freshly updated, and the job always passes while the docs quietly rot. A doc-drift check catches a file that fell behind. It does not tell you whether a new rule does anything. The only way to learn that is to test the rule in isolation. 1. Add or modify the rule 2. Discard the current artifact or stash it on a branch 3. Start a fresh session with the updated rules 4. Re-run the same task 5. Confirm the issue does not recur If you keep the existing artifact and continue, you are still operating in context polluted by the old system. The model may try to reconcile the new rule with the old work instead of applying it cleanly. You cannot tell whether the rule works, or whether you just fixed the symptom by hand. Discard-and-retry is the only clean measurement. Before you add a rule, ask whether the problem is best solved by a line of prose. If a rule can be expressed as a test, a hook, or a permission boundary, write it there instead — those fail loudly when they rot. Reserve the prose for what only prose can carry: Everything else is context the agent can read from the code itself. A rule that restates the code adds tokens without adding signal. An instruction file should end the way a pull request does: with a confirmation step. A short checklist forces the agent to verify before it reports done, and it gives a human something concrete to review in the diff. Before you finish - The command in step 4 matches the one in CI - The path in step 2 still exists on the main branch - You can point to the test that would fail if this rule broke When the file ships with a check, a stale rule stops being silent. It becomes a checkbox nobody can honestly tick, which is exactly the signal a maintainer needs. Here is the shape of the failure, the one that made me stop trusting my own instructions file. The repo had a rule: "Run make check before you finish." A teammate reorganized the Makefile and renamed the target to make ci-check . Nothing referenced the old name, so nothing failed. The agent read the rule, ran make check , got a target-not-found error — and, because the rule said "before you finish," decided the error was a pre-existing environment issue and reported done anyway. It did not update the rule, because the rule did not tell it that it was allowed to. That is the silent-drift failure in one loop. The rule was still on the page, still trusted, and wrong. No test, no linter, no human was in the path to catch it. It was not a discipline failure on anyone's part. The system had no place where staleness was supposed to be detected. A drift check is only useful if it runs. Put it on the same schedule as your dependency updates and your license scans, not as a one-time cleanup you do when you remember. - Every PR that touches a path a doc covers → re-run the drift check - Every model-version upgrade → re-read AGENTS.md and delete rules written for the old model - Every month → a rule you cannot remember triggering is a rule you delete The last one is the hard one. Most rules are added reactively, after a bug. The bug stops happening, the rule stays, and a year later it is a small tax on every session. A rule with no documented rationale and no recent trigger is not insurance, it is rent. Delete it, and let the next real bug re-earn its place. The instinct is to blame the team: someone should have updated the file. But the real issue is that the file had no mechanism to announce its own staleness. A doc with no owner, no covers: line, and no checklist is structurally unable to tell you when it is wrong. You cannot discipline your way around a missing feedback loop. I have not measured the failure rate of stale rules on a large corpus. The covers: -line drift check and the discard-and-retry loop are practices I have seen work in small teams, not numbers I can cite. Treat the mechanism as the argument, and the specific thresholds as a starting point you tune to your repo. What I am confident about is the structure: an instruction file that is trusted but not verified will fail in the direction of the old world. The fix is to give the file the same verification discipline you give the code — a review trigger, and a way to prove a rule still works. This post was written with AI assistance. The author is responsible for its content.