# I Built a Multi-Agent Coding Setup to Keep Me in the Loop

> Source: <https://dev.to/jancera/i-built-a-multi-agent-coding-setup-to-keep-me-in-the-loop-lo3>
> Published: 2026-10-01 21:22:44+00:00

I didn't build this after a big disaster. No agent ever deleted my database or broke a production deploy. The problem was slower and more annoying than that.

Every time I used an AI agent on a real task, I spent too much time managing it. It asked permission for simple things, the implementation often failed when the task was bigger, and even when the code worked, it often ignored the way the project was organized.

Each of these was small, but after a while they made working with AI more tiring than it should be. So I started adding rules, one problem at a time, and the result is [Jancera/skills](https://github.com/Jancera/skills), a set of skills and subagents I use in Google Antigravity.

Most of the talk about coding agents is about taking the human out of the process. I wanted a setup where I stay in it, but without having to watch every step.

Every rule in the repo comes from that. Agents only ask me about decisions that matter, the work is split into pieces small enough for me to review, and nothing reaches the main branch unless I merge it myself.

The agent asked me to approve almost everything, and most of it was commands that only read files. I never cared about approving reads. What I wanted to approve were the things that change something.

Now read-only inspection commands are approved by default. Implementation agents run inside an OS-level sandbox, where commands run without prompts and network access is off unless a task needs it. I only get asked when something goes beyond that, like a commit, a git push, installing a package, downloading a file, or editing outside the project.

When I gave the agent a large feature, the implementation often failed, and I believe the main cause was poor planning. Matt Pocock's talk on breaking requirements into vertical slices changed how I split the work.

The `deepwork` skill breaks a feature into phases, and each phase into vertical tasks. A task touches the whole slice it needs (schema, backend endpoint, UI) and produces something I can run and test. Each task runs in its own git worktree from the same base branch, and tasks are never stacked on top of unmerged work.

When a task finishes, I check out the branch, test it, and merge it myself.

After every task in a phase is merged, an `oracle` subagent reviews the combined diff of the phase for bugs, security issues and regressions. Accepted findings go to a `fixer` agent for mechanical fixes or to a `designer` agent for UI work.

I tested `deepwork` on [shorts-pipeline](https://github.com/Jancera/shorts-pipeline), a side project I work on now and then. I started four tasks at once:

All four branched from the same commit, and each one changed config, the Flask backend, the HTML frontend and the tests within the same task. Together they added about 1,360 lines. The subtitle task alone added 1,068, almost half of them tests.

I started the run and went to do other things, and in about a day all four were ready for review. I didn't reject any of them, and they landed on `main`.

Then the oracle reviewed the merged phase and found problems that no single task had. The shorts-count task and the video-length task both touched the same settings flow, and each one worked alone. After the merge, the settings route was missing validation for the new length field, the stale-state check ignored changes to the shorts count, and some test mocks still used the old function signature.

I didn't catch any of this in my own review, and the fix landed shortly after the last merge. That's the reason the gate looks at the combined diff of the whole phase. When tasks run in parallel, some bugs only exist after everything is merged, and no per-task review will find them.

Without this setup, my guess is that these four tasks would have taken me several days, though I didn't measure it. It also doesn't mean the code is as good as what I would write myself, which I'll come back to below.

Most of the gain came from where my attention went. I was only needed at the gates, and I could walk away in between because nothing reaches `main` without me and the oracle checks what I miss once the tasks come together.

**Project conventions.** The agents tend to ignore the structure that already exists in a project, like its layers and where things belong. The code works, but it doesn't follow the patterns I defined or agree with. The repo has an `explorer` agent that checks module boundaries after a phase, but it's optional and didn't run in the test above, so that's the first thing I'll try.

**The lighter workflow.** `to-spec` and `implement-spec` are meant to be a simpler version of `deepwork` for medium tasks, and they still need work. My hypothesis is that the planning is too shallow, since `implement-spec` doesn't split the work into vertical tasks the way `deepwork` does.

**The sandbox.** I use Antigravity's default sandbox exactly as it comes. I haven't looked into the other sandbox options, or checked whether the default is enough for this workflow, and for now I'm leaving it that way.

Most of this is adapted from other people's work. The `grilling` skill comes from Matt Pocock's grill-me idea, which I picked up from his talk. `deepwork` started from the structure Fabio Akita uses with OpenCode, which I rebuilt on Google Antigravity's own pieces: one worktree per task, OS sandboxing, and subagents with their own model tiers.

The repo is public at [github.com/Jancera/skills](https://github.com/Jancera/skills). If you try it, I'd like to hear where it breaks.
