# I built an agentic workforce where done means a command exited 0

> Source: <https://dev.to/forgecoreai/i-built-an-agentic-workforce-where-done-means-a-command-exited-0-475i>
> Published: 2026-10-11 11:20:03+00:00

Every failure I have had with coding agents looks the same at the end. The agent writes a summary that reads like success: the feature is implemented, the tests pass, the branch is ready. Then I open the thing, and half of it is not there.

The interesting part is that the model is not lying. It has no way to check. It generated the most plausible continuation of the conversation, and "done, with a tidy summary" is an extremely plausible continuation. Nothing in the loop ever asked for evidence.

So I stopped trying to make the model more honest and changed the shape of the loop instead. In the agentic workforce I run now, no model is allowed to say done. A command says done, and it has to exit 0.

When I type an ask, the first thing that happens is not work. It is a contract.

```
/crew add a pricing page with a monthly / annual toggle
```

The card that opens carries six things:

The order is the point. The proof is frozen while the work does not exist yet. At that moment I am uninformed and honest, so I write:

```
npm run test:e2e -- pricing.spec.ts
```

and not "looks right to me". Later, with work already on the branch and the temptation to call it finished, the command is a fact from an earlier conversation. Neither I nor the agent can quietly move it. That is most of the value: not detection, but removing the option.

In the workforce, done is not a status an agent sets. It is the exit code of that command, run as a tool call, at the point where the card would otherwise close.

The model writes the code and the summary. It never writes the verdict.

The second failure mode is drift: three agents on the same branch, each doing something locally reasonable, together producing nothing. So a card has one writer, and a card has one owner - a coordinator that wakes on every card event with the whole record in hand and takes exactly one decision: retry, rescope, split, or ask me.

The asks are bounded too. Before any work starts, the chat asks at most five questions in one batch, or none. After that I am out of the loop until the card ends, and it ends either with a verified result in the thread I asked from, or with one concrete question. Token budgets are per card, so a run that reaches its budget stops instead of burning through the rest of the night, and a worker that fails five tool calls in a row stops itself.

Illustrative card, from the demo board:

```
16:12:04  /crew add a pricing page with a monthly / annual toggle
16:12:09  contract: artifact app/pricing/page.tsx
          proof    npm run test:e2e -- pricing.spec.ts
16:12:11  card opened - coordinator owns it, your turn ends
   ...
16:41:21  done - proof exit 0 - 2 runs, 1 fix
```

That last line is what reaches the chat: not a summary of feelings, the proof that passed and how many attempts it took. If the proof had failed twice, I would have got one question instead, with the two failure lines attached.

Three honest limits, because a post about verification that only lists wins is its own kind of bullshit.

`test:e2e -- pricing.spec.ts` can be green while the toggle
does nothing, if the test asserts the wrong thing. The leverage is all in writing the command well, and that is
human work.
There is also a fourth thing I did not expect: because the proof fixes the artifact, "can you also just..." turns into a second card instead of a scope creep inside the first one. Annoying, and correct.

The generalisable rule is not complicated:

Any agent framework can do this. Most do not, because "done" is much easier to print than to prove.

I built it as a plugin for Hermes Agent - Apache-2.0, self-hosted, runs on your own machine, no telemetry - and the whole thing is [github.com/macd2/Crew](https://github.com/macd2/Crew):

```
hermes plugins install macd2/Crew
python3 ~/.hermes/plugins/crew/install.py --profile NAME
hermes -p NAME plugins doctor crew
```

If you would rather watch it before installing anything - the 53 second launch film, one ask through to the verified result:

[crew.forgecoreai.com](https://crew.forgecoreai.com/) is the same loop as a page, if you would rather read than watch.

So the question I would actually like answered in the comments: **what does done mean in your setup?** Do you have a command, or is it a feeling? I am collecting the ones people genuinely trust - a solo dev with `pytest`, a team with a deploy smoke test, someone checking a number in a dashboard. A proof you cannot write is a task you cannot delegate, and the interesting stories are the tasks where you had to admit that.

*Drafted with an AI assistant and edited by me.*
