A customer asks an agent to change a delivery address. The agent calls a tool, receives a successful response, and says, “Done.” It sounds like an ordinary interaction. But what happened: was a customer profile updated, was a request queued, or did the fulfillment system actually accept the new destination?
Written with AI assistance.
That distinction is the part of agent design I want to examine. My working argument is that a tool result needs to support the promise the agent makes to the customer. An operation can succeed at its own task while leaving the customer’s task unfinished.
Reading the Rails World 2026 session descriptions helped frame this question. Kinsey Durham Grace’s session connects agent behavior with the conventions and implicit knowledge in a codebase [1]. Joël Quenneville’s session emphasizes feedback from tests and review [2]. Taking those ideas into business software, I wonder how much successful use depends on knowledge that the interface never expresses.
Consider a hypothetical address-change service. The customer has been authenticated and authorized to modify the order. The following sequence is a design example, not a production trace.
At 10:00, the customer asks to change the delivery address on order A17. The agent submits a request scoped to that order, rather than changing the default address on the customer profile.
At 10:00:01, the service returns HTTP 202 and request reference R42. Processing is pending. The agent can now say that the request was submitted. It cannot yet say the delivery address changed.
At 10:00:08, the fulfillment service rejects R42 because the parcel has passed its dispatch cutoff. The agent must explain that the requested change could not be completed and offer the next available step.
The important detail is that the first response was not an error. HTTP 202 explicitly means processing has been accepted but remains incomplete; the request may ultimately be disallowed [3]. Turning that response into “Done” would discard information the system already had.
In this example, the failure is in the translation from system state to customer meaning. A more capable model might avoid it. A clearer tool contract would make avoiding it less dependent on the model guessing correctly.
For this service, I would want the agent to receive five things: the affected order, the operation requested, its current state, when that state was observed, and how to check it again. An illustrative response could read: “Order A17; change delivery address; pending; observed at 10:00:01; request R42.” That is a compact description of what the agent knows. It also makes the missing fact visible: there is no confirmed change yet.
The source of confirmation matters. Reading the new address from the customer profile would not prove that fulfillment uses it. Even reading it from the order database may be insufficient if a separate dispatch system has already taken a copy. The product team has to decide which system can establish completion for this particular promise.
This is a limit of the proposal. There may be no single authoritative answer, or the system that has it may be temporarily unavailable. Adding a field called “confirmed” cannot resolve that architectural problem. The agent needs a way to preserve uncertainty instead.
That is where the product question becomes interesting: how much uncertainty should the customer see, and who remains responsible for resolving it after the conversation ends?
Now change one detail. The connection breaks before the agent receives R42. The operation may have executed, or it may never have arrived. Neither success nor failure has been established.
Repeating the operation is safe only under the service’s actual retry contract. Stripe provides a concrete example: its API supports idempotency keys and returns the stored result for subsequent requests using the same key, subject to documented conditions and retention [4]. That behavior must be implemented by the service; inventing a request identifier in a prompt does not provide it.
For the hypothetical address service, I would separate “check the existing request” from “submit another change.” If the service cannot identify the earlier attempt, the agent should route the uncertainty to someone who can investigate. Quietly retrying until something looks successful can obscure what happened. A human handoff is only helpful if it carries that uncertainty forward. The next person needs the attempted action, the available references, and what remains unknown. “Customer wants an address change” leaves out the reason they are being involved.
A useful evaluation would hold the customer’s request constant and vary the service response. These are proposed cases, not measured results.
In the confirmed case, the authoritative system reports the new delivery address for A17. The agent may confirm that specific change, without promising that delivery itself is guaranteed.
In the pending case, the request has been accepted but not completed. The agent should describe it as pending and explain how the result will be communicated.
In the timeout case, the outcome is unknown. The agent should avoid claiming completion or issuing an unsafe duplicate request. If it cannot resolve the state, that unresolved condition should remain visible in the handoff.
I would score the final customer message separately from tool selection. Calling the correct operation and then overstating its result is still a failure. I would also check the opposite error: if completion is established, repeatedly refusing to confirm it creates needless work for the customer.
This approach adds cost. Status checks take time, and some systems cannot offer immediate confirmation. The design therefore needs a stopping rule and an owner for follow-up. “Keep checking” is not a complete workflow either.
None of this makes a stronger model irrelevant. The agent still has to identify the customer’s intent and explain a result clearly. But an interface that hides the difference between accepted and completed asks the model to reconstruct facts it may never receive.
An experienced colleague might already know to ignore a reassuring status and check another screen. That workaround can make a product appear more complete than it is.
When we put an agent in that colleague’s place, we get a chance to examine the workaround. Does the business rule need to be made explicit? Does the interface need a better result? Or is responsibility split between teams in a way that no tool can resolve on its own?
The next time someone says, “The system says it’s done, but you still need to check,” I would ask what that last check establishes. It may be the most useful part of the workflow to understand before automating it.
[1] [Kinsey Durham Grace — Agent-Proof Your Rails App, official session description](https://rubyonrails.org/world/2026/speakers/kinsey-durham-grace).
[2] [Joël Quenneville — Harness Engineering on Rails, official session description](https://rubyonrails.org/world/2026/speakers/joel-quenneville).
[3] [RFC 9110, section 15.3.3 — HTTP 202 Accepted](https://www.rfc-editor.org/rfc/rfc9110.html#name-202-accepted).
[4] [Stripe API documentation — Idempotent requests](https://docs.stripe.com/api/idempotent_requests).
When an AI Agent Says “Done,” What Has Actually Happened? was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.