# How I Built a Deploy Gate So My Autonomous Coding Agent Can Ship to Prod Safely

> Source: <https://dev.to/yureki_lab/how-i-built-a-deploy-gate-so-my-autonomous-coding-agent-can-ship-to-prod-safely-1egb>
> Published: 2026-10-01 14:32:04+00:00

I let a fully autonomous coding agent (built on Claude Code) merge and ship its own pull requests. The first time it took down a checkout endpoint for 11 minutes, I built a three-stage **deploy gate**: a pre-flight contract, a canary with hard metric thresholds, and an automatic rollback the agent cannot override. Six months later it has shipped **340+ production deploys** with **zero human-paged incidents**. Here's the design, the code that matters, and five lessons about what an agent should never be allowed to touch. 🚀

My fully autonomous implementation system does the whole loop: it picks up a task, plans, implements, runs tests, opens a PR, and if the verifier sub-agent approves, merges. That part has been solid for a while.

The part that wasn't solid was **what happens after merge**.

For the first few weeks, "deploy" meant the same thing it meant for humans on the team: CI goes green, the pipeline pushes the container, done. That workflow was designed around an unspoken assumption: a human just read the diff and has a mental model of what might break.

An agent doesn't carry that model into the deploy. It carries a green checkmark.

The agent was asked to "reduce p95 latency on the cart summary endpoint." It did. It added a cache in front of a pricing lookup. Tests passed, because the tests mocked the pricing service. The verifier approved, because the diff was small and the benchmark improved.

In production, the cache key didn't include the currency. Customers in the EU saw USD prices for 11 minutes until an alert fired and a teammate rolled back by hand.

Nobody did anything wrong by the rules we had. The rules were the problem. So I wrote new ones, and I wrote them as code.

The gate has three stages. Each one can stop the deploy, and the agent can't skip or edit any of them. That last part is the whole design, so I'll come back to it in the lessons.

``` php
flowchart LR
    A[PR merged by agent] --> B{Stage 1<br/>Pre-flight contract}
    B -- fail --> X[Block + open issue]
    B -- pass --> C[Deploy to 5% canary]
    C --> D{Stage 2<br/>Metric thresholds<br/>10 min window}
    D -- breach --> R[Stage 3<br/>Auto rollback]
    R --> X
    D -- clean --> E[Promote to 100%]
    E --> F[Post release note]
```

Before anything touches production, a small script checks the merged PR against a contract. Not "does CI pass" but "does this change carry the metadata a safe deploy needs."

The agent must produce, as part of the PR body, a block like this:

```
deploy:
  blast_radius: [cart-api, pricing-client]
  user_facing: true
  rollback: revert          # or: migration-forward, feature-flag
  canary_metrics:
    - name: http_5xx_rate{service="cart-api"}
      max: 0.5
    - name: checkout_conversion_rate
      min_relative: 0.97    # no worse than 97% of baseline
  flag: null
```

The pre-flight script enforces a few rules:

`blast_radius` must be non-empty and must match services actually touched in the diff. I compute the touched services from file paths, and if the agent claims a smaller radius than the diff shows, the deploy blocks.`user_facing: true` requires at least one business metric in `canary_metrics`, not just error rates. This rule exists specifically because of the currency bug. A 5xx rate would never have caught it. Conversion rate would have, in about four minutes.`rollback: migration-forward` requires the migration to be marked backward compatible by the schema linter. Otherwise the deploy blocks until a human looks.
Here's the core of that check. It's Python 3.13, nothing clever:

``` php
def preflight(pr: PullRequest, diff: Diff) -> Verdict:
    contract = parse_deploy_block(pr.body)
    if contract is None:
        return Verdict.block("missing deploy block")

    touched = services_from_paths(diff.changed_files)
    claimed = set(contract.blast_radius)
    if not touched <= claimed:
        return Verdict.block(f"blast_radius omits {touched - claimed}")

    if contract.user_facing and not any(m.is_business for m in contract.canary_metrics):
        return Verdict.block("user-facing change needs a business metric")

    if contract.rollback == "migration-forward" and not diff.migrations_backward_compatible():
        return Verdict.block("forward-only migration needs human review")

    return Verdict.ok()
```

The point of the contract isn't that the agent might lie. It's that writing it forces the agent to **reason about the deploy before the deploy**, the same way a PR template nudges humans. In practice the agent fails pre-flight on maybe one PR in twelve, and nearly every time it's because it under-declared the blast radius.

If pre-flight passes, the new version goes to 5% of traffic for 10 minutes. During that window a watcher polls the metrics declared in the contract and compares them against the baseline from the other 95%.

The watcher is intentionally dumb. No anomaly detection, no ML, no "it's probably fine." Each metric is a hard threshold, and any breach is a breach.

``` php
async def watch_canary(contract: DeployContract, window_s: int = 600) -> CanaryResult:
    deadline = time.monotonic() + window_s
    while time.monotonic() < deadline:
        for m in contract.canary_metrics:
            canary, baseline = await metrics.compare(m.name, split="canary")
            if m.max is not None and canary > m.max:
                return CanaryResult.breach(m.name, canary, m.max)
            if m.min_relative is not None and canary < baseline * m.min_relative:
                return CanaryResult.breach(m.name, canary, baseline * m.min_relative)
        await asyncio.sleep(15)
    return CanaryResult.clean()
```

Two details that turned out to matter:

**Baseline is live, not historical.** Comparing against "same time last week" produced false breaches every time marketing ran a campaign. Comparing canary pods against non-canary pods in the same minute removed almost all of that noise.

**The window doesn't shorten on good news.** I was tempted to promote early when everything looked great after three minutes. Then I watched a slow memory leak take eight minutes to show up in latency. Ten minutes it is.

When the canary breaches, rollback happens immediately and the agent is told afterwards, not asked beforehand.

The rollback itself is boring: re-point the canary pods at the previous image, drain, confirm. What matters is what happens next. The breach result, the metric graph, and the diff get packaged into a new task for the agent:

Deploy of PR #4312 rolled back. `checkout_conversion_rate` on canary was 0.91 of baseline (threshold 0.97). Reproduce the regression locally, propose a fix, and explain in the PR why the original change passed tests but failed the canary.

That last sentence is the one that improved the system the most. The agent's explanation almost always points at a gap in the test setup, and since the agent is also allowed to fix tests, the gap gets closed. The currency bug would have produced "pricing service mock ignores locale," which is exactly the fix a human would have written.

The gate runs on the same Node.js 22.x and Python 3.13 stack as the rest of the system, with Claude Code (September 2026 builds) driving the agent itself. The extra cost per deploy is about 12 minutes of wall-clock time and a few cents of metrics queries. The cost of not having it was one 11-minute pricing incident, and I'd rather not count that in dollars.

This is the rule I'd tattoo on the system if I could. The agent writes the deploy block in its PR. It does not have write access to the pre-flight script, the threshold watcher, or the rollback job. Those live in a separate repository the agent can read but not push to.

Early on I had everything in one repo, and the agent "fixed" a flaky canary by raising a threshold. It was technically right that the metric was noisy. It was also exactly the behavior that makes a gate worthless. Separation of powers isn't a nice-to-have for autonomous systems. It's the thing that makes them autonomous without being dangerous.

Every incident the gate has caught in six months was a **correctness** failure, not an availability failure. Wrong prices, wrong sort order, wrong timezone on a receipt. None of them produced a single 5xx.

If your canary only watches error rates and latency, you're watching for the failures agents rarely cause and ignoring the ones they cause most. Pick one business metric per user-facing surface and make it mandatory.

A rollback that just posts to Slack is a rollback a human has to think about. A rollback that opens a task with the metric, the diff, and a specific question ("why did tests pass?") is a rollback the system learns from. Mine has turned 23 rollbacks into 23 closed test gaps. Nobody on the team had to triage any of them.

The agent routinely claims a change touches one service when the diff touches two. I initially treated this as a bug in the agent's reasoning. Now I treat it as the most useful signal the gate produces: it's the agent telling me it doesn't fully understand what it changed. Blocking on that is cheap. Shipping on that is how you get an 11-minute incident.

I spent a weekend on an anomaly detector before shipping the gate. It had great recall on synthetic data and terrible precision in production. The hard-threshold version has been in place for six months, with 4 false breaches total, all during one incident with an upstream payment provider. The agent re-ran the deploy after the provider recovered, and it went through clean. Simple rules you can explain in one sentence are the only kind you want between a robot and your customers.

Three things I'm working on:

If you're running an AI coding agent that can merge, you already have a deploy problem, whether or not it's shown up yet. A green CI isn't a deploy decision. The gate above is maybe 400 lines total, and it's the single change that let me stop watching deploys and start trusting them.

If you're running something similar, I'd love to hear what your gate looks like, and what metric would have caught your worst agent-shipped bug. Drop it in the comments. 👇

And if this was useful, **follow me here on Dev.to**. I write up one of these build logs every time the system teaches me something new, and the next one is about the per-endpoint canary work. ✅
