When you hand an agent work with external effects, two worries always come up.
First: when the process dies, what happens to the work in progress? An agent d at an approval gate crashes β can it resume, or does everything restart from scratch?
Second: when an LLM retries, does it execute the same side effect twice? Retrying a send or a data mutation carries a real double-execution risk.
My previous article measured HITL approval, audit trails, and structured output across three frameworks. It recorded a failure where Strands double-fired a publish call with the identical draft. This article measures what comes next: 34 more runs, same task, same model, same recorder proxy.
Repo (all code, traces, and analysis scripts open):
https://github.com/sunnydachs/agent-framework-showdown
(Read the previous article here. here.)
Same task, same model, same recorder proxy: 3 frameworks x 3 cells x multiple runs.
[agent] β [approval gate] β [publish (destructive)]
β β
crash it here s waiting
The three cells:
publish under 3 key strategies (position key / content hash / no key) and count duplicate executions on retry
The three idempotency key strategies:
position key: {workflow}:{step}:{tool}
content hash: sha256(the raw arguments)
no key: nothing
SIGKILL the agent while it waits for approval, then resume in a new process:
| Framework | Persistence | Resume time | State survived | LLM calls to resume |
|---|---|---|---|---|
| LangGraph (durable checkpointer) | checkpointer on disk | 0.01-0.02s | 3/3 | zero |
| LangGraph (no checkpointer) | none | - | 0/3 | - |
| Strands | none built-in | 4.9s avg | full re-run | 5.3 avg |
| CrewAI | none for agents | 4.2s avg | full re-run | 2.0 avg |
LangGraph's durable checkpointer persists the graph state to disk even while the process is dead. The new process restores to that position in 0.01s β with zero LLM calls. The difference between resume and redo is only whether the state lives outside the process.
Without a checkpointer, an identical-looking "resume" is a full re-run: Strands runs its average 5.3 LLM calls again and pays the full token cost a second time.
LangGraph has a known issue here (#8764): if the process dies before the first checkpoint is persisted, recovery may find no checkpoint and no record that the run was ever accepted. On the version I tested, resuming the empty thread succeeded without raising β so the behavior is version-dependent. Don't rely on the error either way; keep an external acceptance ledger.
Average duplicate executions when an LLM retry re-calls publish:
| Key strategy | Same-args retry | Reworded retry |
|---|---|---|
| Position key | 1 dup, all deduped | 3 dups, 33% deduped + rest rejected as caller bug |
| Content hash | 1 dup, all deduped | 1 dup, 0% deduped β it slipped through |
| No key | 1.33 dups, 0% deduped | 1 dup, 0% deduped |
This is the core result. The content-hash key silently fails the moment the model rewords the arguments on retry.
The reason is simple. An LLM retry does not replay the saved HTTP request. It reasons again from a context that now includes the timeout error, and emits a new tool call. The arguments get reworded, the order changes, fields appear. With sha256(args) as the key, the retry produces a different hash, sails past the dedup check, and executes the side effect a second time.
The position key ({workflow}:{step}:{tool}) identifies the intent β where the call sits in the workflow β not the bytes. The same position with the same operation yields the same key no matter how the arguments change.
One more measurement: every duplicated publish carried a different tool_call ID on the wire. Nothing at the protocol layer detects "this is the second time for this operation." Detection lives at the recording level only.
Note: the position-key strategy rejects "same position, different arguments" calls as a caller bug (2.67 of the runs here). That is by design β it flags intent drift instead of letting the key be reused.
Can an auditor reading the traces alone recover these four facts:
But this is only because the recording is at the wire level. The proxy keeps every attempt β first call, dedup, caller-bug rejection β as its own record, so the auditor can reconstruct everything.
Framework-level trace surfaces show none of the double-firing. What matters in audit design is where the evidence lives.
Across the 34 runs, each key strategy fails differently:
For enterprise use, the silent failure is the scariest class. The intuition "a hash key makes retries safe" breaks the moment the caller becomes non-deterministic β that is the conclusion from these measurements.
Everything is open. The README has the commands for all experiments (34 runs here + 72 from the earlier articles):
This is a personal OSS project β no warranty. Use at your own risk, and issues are welcome.