My fully autonomous implementation system kept making the same handful of mistakes every single week, and I was the one patching its instructions at 11pm. So I built a post-mortem loop: every failed run produces a structured incident file, a separate agent distills those into one-line rules, and each rule has to pass a replay regression check before it's allowed into the agent's playbook. Over three months, repeat failures dropped from 38% of all failures to 6%. Here's the design, the load-bearing code, and what I'd do differently. 🚀
Some context first. I run a 24/7 autonomous dev system built on Claude Code (the 2.x line as of this writing). An orchestrator module picks tasks, parallel implementation agents do the work in isolated worktrees, a self-healing agent retries broken builds, and I review the output in the morning.
It works. Most days I wake up to a few mergeable PRs. But after about two months in production I did something I should have done on day one: I classified every failure from the previous twelve weeks.
| Failure bucket | Share |
|---|---|
| Genuinely new problem | 41% |
| Flaky infra / API outage | 21% |
| Something we had already seen before | 38% |
That 38% bucket hurt. Same mistakes, different Tuesday:
packages/api, then "fixing" the resulting import errors
Each repeat cost me roughly 25 minutes of review and cleanup, plus a few dollars of tokens. Worse, it eroded trust. An agent that makes new mistakes is learning. An agent that makes the same mistakes is a liability.
And here's the embarrassing part: I had a lessons-learned document. It was 900 lines long. The agent either never read it, or read it and treated it as background noise. I was writing lessons for a reader that didn't exist.
The constraint that made this interesting: I didn't want to add a human to the loop. The whole point of the system is that it runs without me. So the fix had to be a mechanism, not a habit.
The loop has three stages: capture, distill, admit. Each stage is owned by a different agent with a different prompt, and each stage produces a file the next stage reads.
flowchart LR
A[Failed run] --> B[Capture: incident file]
B --> C[Distill: proposed rule + replay scenario]
C --> D{Admit gate: replay passes?}
D -- yes --> E[Playbook]
D -- no --> F[Rejected, logged]
E --> G[Next run reads playbook]
G --> A
When a run fails for any reason (test failure, watchdog kill, human rejection in review), the orchestrator asks the agent that just failed to fill in a fixed template before the context is thrown away. The template is small on purpose:
task: "Add rate limiting to /v1/upload"
failure_class: wrong_assumption # wrong_assumption | bad_tool_use | scope_creep | env | unknown
trigger: "CI failed: 14 import errors after agent moved tests"
believed: "Tests are run from repo root with `pytest`"
actual: "Each package has its own pytest.ini; must cd into packages/api first"
cost_minutes: 22
resolved_by: human
The two fields that matter are believed and actual. A stack trace tells you what broke. The delta between what the agent believed and what was true tells you why, and that delta is the only thing worth turning into a rule.
I tried free-form post-mortems first. They were long, apologetic, and useless. Forcing a one-line believed and a one-line actual made the agent actually commit to a diagnosis.
Every ten incidents (or weekly, whichever comes first), a distiller agent reads all unresolved incident files and proposes rules. It is not the agent that failed. That separation turned out to matter a lot (see lesson 4).
The distiller's prompt has three hard constraints:
provisional.
Here's what a proposed rule looks like coming out of the distiller:
rule: >
Before running any test command, check for a package-local test config
(pytest.ini, jest.config.*, vitest.config.*) and run from that directory.
cites: [2026-07-14-0932, 2026-07-21-1105, 2026-08-02-0847]
status: candidate
replay:
repo: fixtures/monorepo-two-packages
task: "Fix the failing test in packages/api/tests/test_limits.py"
pass_if: "agent runs pytest with cwd == packages/api"
Notice the rule is boring. That's the goal. Clever rules get ignored; boring, specific, checkable rules get followed.
This is the stage that made the difference. A rule does not enter the playbook because it sounds reasonable. It enters because it demonstrably fixes the replay and doesn't break anything else.
The admit gate runs the agent on the replay scenario twice: once with the current playbook, once with the current playbook plus the candidate rule. Then it runs the full replay suite (about 30 scenarios by month three) with the candidate included.
from dataclasses import dataclass
@dataclass
class ReplayResult:
scenario_id: str
passed: bool
cost_usd: float
def admit(candidate: Rule, playbook: Playbook, suite: list[Scenario]) -> bool:
before = run_replay(candidate.replay, playbook)
after = run_replay(candidate.replay, playbook.with_rule(candidate))
if before.passed or not after.passed:
reject(candidate, reason="replay not discriminative")
return False
regressions = [
r for r in run_suite(suite, playbook.with_rule(candidate))
if not r.passed
]
if regressions:
reject(candidate, reason=f"regressed {[r.scenario_id for r in regressions]}")
return False
if len(playbook) >= 60:
playbook.retire(playbook.coldest())
playbook.add(candidate)
return True
Two details worth calling out:
last_hit timestamp (updated whenever the agent explicitly cites it in a run). When the cap is reached, the coldest rule gets retired to an archive. This single constraint is why the playbook stayed readable instead of becoming another 900-line document.
| Metric | Month 1 | Month 3 |
|---|---|---|
| Total failed runs | 71 | 64 |
| Repeat failures (share) | 38% | 6% |
| Rules admitted | 9 | 44 (cumulative) |
| Rules rejected by gate | 4 | 19 (cumulative) |
| Rules retired (cold) | 0 | 7 |
| Avg replay suite cost per admit | $1.10 | $3.40 |
Total failures barely moved, which surprised me at first. The loop doesn't stop the agent from hitting new problems. It stops it from hitting the same problem twice. That's exactly what I wanted: new failures are information, repeat failures are waste.
The replay suite cost per admission tripled because the suite grew. I'm fine with that. Three dollars to permanently retire a 25-minute recurring mistake is the best trade in the whole system.
believed vs. actual was the single highest-leverage design decision. Logs tell you what happened. The delta tells you what to change. If your post-mortem template doesn't force the agent to state what it assumed, you're collecting noise.
Before the admit gate, I had rules like "be careful with monorepos." Totally unfalsifiable. Requiring every rule to ship with a replay scenario filtered out everything vague, and the ones that survived were specific enough that the agent could actually act on them.
Instruction files for AI agents decay the same way wiki pages do. Nobody prunes them, so they grow until they're ignored. A hard cap plus a "retire the coldest" rule keeps the playbook under what the agent can genuinely hold in attention. Sixty felt right for my system. Yours may be forty.
Early on, I had the failing agent propose its own rule in the same context. The rules were defensive and over-specific ("never move test files"). A fresh distiller agent reading three incidents at once wrote better, more general rules because it wasn't trying to justify itself. Separation of concerns applies to agents too.
About a third of admitted rules turned out to be patching ambiguity in my own task specs. "Fix the failing test" with no mention of which package. The loop made my sloppiness visible and measurable, which was uncomfortable and extremely useful.
If your coding agent makes the same mistake twice, that's not a model problem. It's a feedback loop problem. Capture the belief delta, make rules replayable, gate admission, cap the playbook. That's the whole trick.
I'd love to hear what your agent's most repeated mistake is. Drop it in the comments, and if you've found a better admit criterion than "the before-run must fail," I genuinely want to steal it. 💡
Follow me here on Dev.to for more build logs from running a fully autonomous implementation system in production. Next up: how the replay suite itself gets maintained without turning into a second codebase. ✅