How to Stop Two AI Agents from Doing the Same Job Twice A technical guide published on webhook reliability outlines four steps for preventing two AI agents from performing the same business task twice: give the task a stable business identity independent of the worker, claim ownership atomically via a database transaction or conditional update, fence stale workers with ownership generations, and reconcile uncertain external actions before repeating them. The guide warns that a separate read followed by a write invites a race, that expiry does not stop a paused process from resuming, and that a lock cannot by itself guarantee an external email, booking or payment happens exactly once. It cites the PostgreSQL explicit-locking documentation for row locks and application-defined advisory locks, noting neither category automatically describes the lifecycle of a long agent task. Stop duplicate agent work by giving the business task a stable identity, claiming it atomically and checking ownership again where changes are accepted. Use expiry so abandoned work can recover, but do not confuse expiry with cancellation of the old worker. A lock can coordinate access to a local resource; it cannot by itself guarantee that an external email, booking or payment happens exactly once. 1. 01Identify the business taskRetries must refer to the same task rather than create unrelated work. 2. 02Claim atomicallyTwo workers must not both succeed at checking and taking ownership. 3. 03Fence stale workersExpiry does not stop a paused process from resuming later. 4. 04Reconcile uncertain actionsA timeout may hide a completed external action, so inspect before repeating it. 01 — Business identityName the task independently of the worker Duplicate work often starts before an agent runs. The same request may arrive through a repeated webhook, a user retry or two queues that describe the same outcome. If each arrival gets a new unrelated task identity, the workers can behave correctly within their own records and still perform the business action twice. Choose a stable key based on the intended business operation and its scope. A hypothetical shipment-notification task might refer to a shipment identifier and notification version, rather than the worker's run identifier. A later corrected notification needs a distinct version; a retry of the original should preserve the original identity. The key design expresses which operations the business considers equivalent. Keep worker attempts separate from the task. Attempts can have their own identifiers, logs and timing without creating new business outcomes. Our webhook reliability reference https://www.digitalapplied.com/blog/webhook-reliability-idempotency-retries-engineering-reference-2026 explains event-delivery retries; this guide concerns multiple workers competing to fulfill the same task after those events arrive. Test the key design against repeated legitimate work as well as duplicates. Two distinct shipments may need similar notifications, and one shipment may need a later correction. A key that is too broad suppresses valid work; a key that is too narrow permits duplicate effects. Define equivalence in business terms and include the relevant version or operation scope. Do not derive it from an arbitrary model-generated task title that can change between retries. Task identity Names the outcome that should not be duplicated. Worker attempt Tracks which process tried the job and what it observed. Ownership generation Lets the destination reject work from a superseded owner. 02 — Ownership claimMake the claim a single atomic decision A separate read followed by a write is an invitation to a race. Two workers can both read that a task is available before either records ownership. Use a database transaction or conditional update that combines the eligibility test and ownership change, and treat the returned result as the evidence that a worker acquired the task. The PostgreSQL explicit-locking documentation https://www.postgresql.org/docs/17/explicit-locking.html describes row locks and application-defined advisory locks. These are coordination tools with different lifetimes and semantics. Advisory locks depend on participating applications using them correctly, and transaction-level locks end with the transaction. Neither category automatically describes the lifecycle of a long agent task. Keep the claim transaction short. Do not hold a database transaction open while a model thinks or a person responds. Record the durable task state, owner and generation, then perform the long work under an explicit ownership policy. A claim that exists only in a process's memory disappears precisely when recovery needs it most. Make the losing claim path explicit. A worker that fails to acquire ownership should observe or exit according to policy, not continue because its task seems urgent. If it reports status to a user, it should describe the existing task rather than announce that a second attempt has begun. This also prevents duplicate planning and repeated tool lookups from consuming resources even when the final write happens to be protected later. The model should not decide whether it owns the task from a conversational message. The durable claim result is the authority for starting work. 03 — Lease lifecycleUse expiry without mistaking it for cancellation A lease gives a worker ownership for a bounded interval and allows recovery if it stops renewing. Store the deadline in the authoritative system and renew it through a conditional operation tied to the current owner and generation. A renewal must fail after ownership has moved, rather than allowing an old worker to extend a new worker's claim. A paused worker may resume after its lease expires. It can still have network access, a pending tool result and a plan to continue. Expiry changes the coordinator's view of authority; it does not physically stop the process. That distinction is why a lease alone cannot prevent duplicate external actions. Choose the renewal and expiry policy around the job's expected duration and failure behavior. Too short an interval creates unnecessary takeovers during ordinary delays; too long an interval delays recovery. There is no universal correct duration. Record lease-loss as a state the worker must handle, and check ownership before progressing to a consequential action. Renewal success should come from the authoritative store, not from a local timer firing. If the worker loses connectivity and cannot establish that the lease remains valid, it must not assume it still owns the task. That uncertainty is especially important before an irreversible action. A local clock can help decide when to attempt renewal, but it cannot override a failed conditional update or a newer owner recorded by the coordinator. - Use authoritative time for claim and renewal decisions. - Require the current owner and generation on every renewal. - Treat lease loss as a reason to stop initiating new work. 04 — Fencing checkReject stale work where changes are accepted A fencing value is a monotonically advancing ownership generation checked by the resource accepting a change. When ownership moves from generation A to generation B, work carrying the older generation must no longer be accepted. The check matters at the mutation boundary, not merely inside the agent's prompt or at the start of a long task. In a hypothetical timeline, worker A acquires a task and pauses. Its lease expires; worker B acquires a newer generation and begins recovery. When A resumes, it tries to write a result with the old generation. A destination that checks the generation rejects that write. Without this destination-side check, the old worker may continue even though the coordinator has correctly reassigned the task. Fencing is only as broad as the resources that enforce it. A local database can condition an update on the current generation, but an external API may not understand that value. A gateway can reject stale submissions before dispatch, yet it cannot retract a request already accepted downstream. External effects therefore need their own duplicate-control and reconciliation strategy. Consider the gap between checking ownership and sending the action. If worker A checks its generation, pauses and then submits after B has taken over, a preflight check alone is insufficient. The accepting resource must validate the generation with the change, or the design must rely on destination idempotency and reconciliation for that external operation. This is why the location of the check matters as much as the existence of a token field. A fencing token that is logged but never checked at the accepting resource provides no exclusion guarantee. An in-flight external request remains a separate recovery problem. 05 — Side-effect controlGive external actions a stable operation key When the destination supports idempotency, send a stable operation key for the intended action across retries and replacement workers. Confirm the endpoint's documented behavior, retention window and treatment of changed payloads. A new random key on every attempt defeats duplicate suppression even if each request includes an idempotency field. Record the request intent before dispatch and the destination's result afterward. If the call times out, mark the outcome uncertain rather than failed. The destination may have completed the action while the response was lost. A replacement worker should query the operation or resulting state when possible before deciding whether another request is appropriate. If the destination has no suitable idempotency or lookup mechanism, acknowledge the remaining limitation. A durable outbox can coordinate recording an intent with local state, but it does not magically make an arbitrary external side effect exactly once. High-consequence ambiguous outcomes may need human reconciliation. The bulk-job engineering guide https://www.digitalapplied.com/blog/bulk-llm-job-engineering-batching-idempotency-qa-2026 covers related retry mechanics at batch scale. Treat payload changes as a separate decision. Reusing an operation key with a different amount, recipient or booking detail may be rejected or behave according to destination-specific rules. Preserve the original intent and define whether a correction is an update to an existing operation or a new operation after reconciliation. An agent should not resolve an idempotency conflict by silently generating a fresh key and repeating a potentially completed action. | Coordination boundaries for the proposed design. Guarantees depend on the actual database and destination implementation. | | | |---|---|---| | Control | What it helps with | What it does not prove | |---|---|---| | Atomic claim | One successful current claim decision | No future stale worker | | Lease | Recovery from abandoned ownership | Cancellation of a paused process | | Fencing | Rejection of stale generations at checked resources | Retraction of an external in-flight request | | Idempotency key | Duplicate suppression under destination rules | Unlimited exactly-once behavior everywhere | 06 — Recovery decisionRecover from the last observed business state A replacement worker should inspect the task record and destination evidence before starting the plan again. Distinguish work that was only proposed, work prepared for dispatch, a request with an uncertain outcome and an action confirmed at the destination. These states require different next steps even if all occurred during a run that eventually failed. For a hypothetical booking workflow, a lost response after submission should lead to a booking lookup using the recorded operation reference. If the booking exists, the worker records the result and continues from it. If the destination definitively rejected the request, a corrected attempt may be appropriate. If the outcome cannot be determined, the system should surface that uncertainty instead of inventing completion or blindly resubmitting. Keep recovery authority narrow. A task owner may be allowed to inspect and reconcile an operation without being allowed to cancel it or issue a replacement. Our handoff ownership guide https://www.digitalapplied.com/blog/ai-agent-handoff-work-ownership separates responsibility for the outcome from permission to perform every action. The same distinction applies when ownership transfers automatically. A completed result should include the destination identifier or other verifiable evidence, not just a worker's final message. That evidence lets a replacement distinguish successful external work from an optimistic local status. If the destination state can later change independently, record when it was observed and avoid presenting the old observation as current indefinitely. Recovery is a sequence of evidence-based transitions, not a replay of the previous worker's narrative. - Read the last durable intent and destination result. - Resolve uncertain external outcomes before retrying. - Preserve the original business operation key during recovery. 07 — Failure rehearsalExercise pauses and races before deployment A useful test deliberately pauses the first worker after claim, after intent recording and after external submission. Let its lease expire, allow a second worker to acquire the task and then resume the first. Inspect which writes are accepted, which are rejected and whether the external action is duplicated. Ordinary successful runs rarely exercise the dangerous ordering. Also test simultaneous claims, failed renewals, a destination timeout and a crash after success but before local completion is recorded. Define the expected invariant for each case: one current owner, stale writes rejected where fenced, and no unexamined repeat of an uncertain action. These are proposed test scenarios, not claims that a particular production system has passed them. Keep the test destination isolated and inspect its state directly. Logs that say only one agent reported success are insufficient if two bookings were created. Our AI transformation service https://www.digitalapplied.com/services/ai-transformation uses observable acceptance criteria so concurrency tests verify the business effect rather than the agents' confidence about their own work. Use a deterministic clock in a small coordination test so expiry can be advanced deliberately instead of waiting for wall time. Assert that only one competing claim succeeds and that a stale generation cannot update the protected result after takeover. Then test the real integration separately with delayed and lost responses. A local state-machine test can establish its own transition behavior, but it cannot establish how a remote service handles an in-flight request. The worked timeline is a design example. Test it against the real database, transport and destination semantics before treating any coordination guarantee as established. 08 — Operational evidenceMake duplicate prevention visible to operators Operators need to see the task identity, current owner, ownership generation, lease deadline and latest action state without reading a full model transcript. They also need a clear distinction between waiting, actively owned, completed and uncertain. A task that has stopped renewing should not remain displayed as normal progress indefinitely. Track rejected stale attempts and repeated claims as diagnostic events. They can indicate normal recovery, an overly aggressive lease policy or a broken worker that continues after losing authority. Do not hide them merely because the final task completed; the pattern can reveal a failure that will matter more on another destination. The goal is controlled completion, not the fiction that work is attempted only once. Distributed systems retry, workers pause and responses get lost. A robust agent workflow makes those events recoverable by preserving task identity, checking current authority and refusing to confuse an unknown external outcome with permission to repeat the action. Keep administrative recovery actions explicit too. An operator may release a stuck task, transfer ownership or mark an external action reconciled, but each change should preserve who made it and the evidence used. Clearing a lock without inspecting the action state can cause the same duplication as an automatic retry. The dashboard should guide the operator toward reconciliation, not offer a generic reset button that erases the history needed to recover correctly. - Show ownership and external-action state separately. - Retain stale-attempt evidence for diagnosis. - Give uncertain outcomes a named recovery owner. Coordinate the action, not just the agents Start with a stable business-task identity and an atomic claim. Add expiry for recovery, destination-side fencing where possible and idempotency under the external service’s documented rules. When the outcome is unknown, reconcile it before repeating it. That is the boundary a lock alone cannot provide.