Most agent frameworks model a step as "do some work, then call the tools."
So the tool call β send the email, call the webhook, charge the card β happens
inside the run, somewhere in the middle of your graph.
That works until it doesn't. The moment you add a retry, a crash-resume, or a
deterministic replay, the same step runs again, and the side effect fires again.
State you can version and reproduce; the world you can't.
This post is about the outbox I shipped in
reactifact 0.14.0 β still the design in
the current 0.15.x line. The produce records what should happen as an
artifact in the context (committed atomically with everything else); the runtime
delivers it after the commit, once per stable id. Replay reconstructs the
answer without re-sending anything.
TL;DR β Don't do I/O inside an agent step. Record the intent as state
(committed atomically with everything else) and let the runtime deliver it
once, after the commit. Retry, resume and replay then rebuild the decision
without re-sending.
Runnable in about a minute, no API key:
pip install reactifact
python -m examples.outbox.main # seven cases, each one asserted
Code: github.com/bzdvdn/reactifact
If you know the graph frameworks, the difference is where the effect lives:
| Typical agent step | reactifact | |
|---|---|---|
| The step does | calls the tool mid-run | writes the intent as state |
| On retry / resume | fires the side effect again | re-derives the same stable id β no second send |
| On replay | re-executes β re-sends | reconstructs the record β never re-sends |
| On parallel/merge | two commits, two sends | one intent, one delivery |
Here's the shape that causes trouble. A produce reacts to an order and sends a
confirmation email:
async def process_order(order):
email.send(order.customer, "Your order is confirmed") # side effect, mid-run
return Receipt(order_id=order.id)
Now consider the three things every long-lived agent eventually needs:
You can try to guard it. A create_once(...)-style check helps for state, but the
guard resolves at commit time, and two producers in one generation share the
same pre-commit snapshot β so both pass the guard and both send. The check is in
the wrong place: it's protecting the write, while the side effect happens
before the write.
The outbox inverts the order:
PendingAction β and
does In reactifact that reads like this:
from reactifact import PendingAction, ProduceCall, produce
@produce(Receipt, also_creates=[PendingAction])
async def process_order(call: ProduceCall) -> None:
order = call.trigger
call.effects.act(
"notify",
key=f"notify:{order.data.id}",
payload={"to": order.data.customer, "order": order.data.id},
)
call.effects.upsert(
Receipt(order_id=order.data.id, text="confirmed"),
id=f"receipt:{order.data.id}",
)
return None # nothing is applied until the runtime compiles the effects
effects.act(...) creates a PendingAction under the stable id
action:{key} and returns None if one already exists (an idempotent re-run).
Note what the produce does not do: it never touches the network. It only states
a change.
Delivery is a separate, injected step:
async def dispatch(context, action):
if action.data.idempotency_key in sent:
return
sent.append(action.data.idempotency_key)
runtime = Runtime(ctx, agents=[Notifier()], dispatcher=dispatch)
await runtime.arun()
The runtime drains the outbox after each generation's commit, marking each action
dispatched β or failed (and re-raising) if the dispatcher throws. A failing
dispatch is state, not a lost effect.
Context effects.act returns None and no second intent exists.dispatched side wins regardless of which branch is the merge target β so
a merge can never resurrect an already-sent action. Divergent payloads under
the same id are an explicit merge You get the operational story you'd want from a queue β but the "queue" is just
versioned state you already have.
This is deliberately not a message broker, and the limitations are worth stating
plainly:
idempotency_key you pass to the
external system, which dedupes on its side. That is the same contract every
real outbox relies on.arun() or an explicit
flush_pending_actions() drains the outbox; the framework doesn't run a poller
for you. Retry and backoff policy stay with the application (wrap your
dispatcher), because that's a product decision, not a framework reflex.
That's the point of the split: state is versioned and reproducible, the world
isn't, and the boundary between them should be something you declare rather
than something buried in the middle of a function.
The whole thing ships as an executable, no-API-key demo. examples/outbox walks
seven cases and asserts each one:
.venv/bin/python -m examples.outbox.main
php
1. commit -> dispatch sent=['notify:42'] status=dispatched
2. re-derivation first run sent 1; the re-run sent 0
3. same generation producers=2 intents=1 sent=1
4. two merged branches branches=2 intents-after-merge=1 sent=1
5. replay sent before replay=1 after replay=1
6. failure -> retry raised 'smtp temporarily unavailable'; retry sent once
7. app-owned retry transient outage absorbed by the wrapper
It shipped in 0.14.0, alongside correlated structured logging and a hard
per-turn deadline with graceful shutdown (the current release is 0.15.x):
docs/en/patterns.md#outbox-external-side-effects and
docs/en/durability.md#outbox-in-production
That's half the story: the outbox keeps an external effect from firing twice.
The other half β making the state behind an answer inspectable and
reproducible, so you can hash it, replay it and audit the trace β is the
If you're building agents that touch the outside world, I'd genuinely like to
know: where do you draw the line between a replayable computation and an external effect? Intents-as-state is one answer β I'm curious about others.