{"slug": "your-ai-execution-finished-the-outcome-hasn-t-happened-yet", "title": "Your AI Execution Finished. The Outcome Hasn't Happened Yet", "summary": "A developer argues that AI coding agents' terminal execution states like COMPLETED or SUCCESS describe only the runtime's own lifecycle, not the eventual real-world outcome, since evidence such as a pull request being merged can arrive minutes later from another system. The writeup proposes separating the execution timeline from the outcome-evidence timeline, so that execution.status = COMPLETED can coexist with outcome.status = UNKNOWN without implying failure.", "body_md": "An AI coding agent finishes its work at 10:00.\n\nIt generated the requested change, returned the result and reached the terminal state we expected.\n\n```\n10:00\n\nexecution_123\n  status: COMPLETED\n```\n\nFrom the execution's point of view, there may be nothing left to do.\n\nBut imagine that the result is a pull request.\n\nAt 10:02, the automated tests finish successfully. At 10:08, a human reviewer approves the change. At 10:17, the pull request is finally merged.\n\n```\n10:00  agent completed\n   │\n10:02  tests passed\n   │\n10:08  human approved\n   │\n10:17  PR merged\n```\n\nNow look again at what we knew at 10:00.\n\nThe execution had finished. The event that would later tell us the pull request was actually merged had not happened yet.\n\nThat creates an architectural problem I find more interesting than simply asking whether `COMPLETED` means `SUCCESS`.\n\n**What happens when the evidence we need to evaluate an outcome belongs to another system — and arrives after the execution itself is already over?**\n\nI've been exploring this question while thinking about the economic evidence needed underneath AI monetization infrastructure.\n\nThe useful distinction, at least for me, is not that workflow completion is somehow a weak or incorrect state. `COMPLETED` may describe the execution perfectly.\n\nThe problem appears when we expect that execution state to also tell us something that can only become knowable later.\n\nA workflow state machine often has a shape like this:\n\n```\nPENDING\n   ↓\nRUNNING\n   ↓\nCOMPLETED\n```\n\nOr perhaps:\n\n```\n      ┌── FAILED\n      │\nRUNNING\n      │\n      └── COMPLETED\n```\n\nThose states answer an important question:\n\nWhat happened to the execution?\n\nDid it start? Is it still running? Did it terminate? Did the runtime encounter a failure?\n\nNothing about that model is inherently incomplete.\n\nThe problem begins when another question gets attached to the same terminal state:\n\nDid the intended outcome happen?\n\nFor our coding agent, those questions can exist on different timelines.\n\n```\nEXECUTION\n\n10:00\nREQUEST\n   ↓\nRUNNING\n   ↓\nCOMPLETED\n   │\n   │   outcome not yet known\n   │\n   ▼\n10:17\nPR MERGED\n```\n\nAt 10:00, the runtime cannot observe an event that will not occur until 10:17.\n\nThat sounds obvious when written as a timeline. It becomes less obvious when the data model compresses both ideas into something like:\n\n```\nstatus = SUCCESS\n```\n\nWhat does `SUCCESS` mean here?\n\nDid the agent successfully finish? Did the generated change pass validation? Was the artifact accepted? Was the pull request merged?\n\nThese can all be legitimate success conditions. They just don't describe the same fact.\n\nA product may intentionally define execution success as successful technical completion. There is nothing wrong with that.\n\nThe ambiguity appears when the same state is later reused as evidence that a different domain condition was satisfied.\n\nThis is the mental model I've found more useful:\n\n```\nEXECUTION TIMELINE\n        ≠\nOUTCOME-EVIDENCE TIMELINE\n```\n\nI'm using those terms descriptively rather than as a formal distributed-systems taxonomy.\n\nThe point is simply that execution state and evidence about the resulting outcome may evolve independently.\n\nThat means this state is not contradictory:\n\n```\nexecution.status = COMPLETED\noutcome.status   = UNKNOWN\n```\n\n`UNKNOWN` does not mean the outcome failed.\n\nIt means the system does not yet have enough evidence to make the outcome claim we're asking it to make.\n\n**Not yet known is not the same as failed.**\n\nAnd once that evidence can arrive later, the problem stops being only about workflow state.\n\nIt becomes a distributed-systems problem.\n\nGo back to the coding agent.\n\nAt 10:00, its execution completes and produces a pull request.\n\n```\nexecution_123\n      ↓\nartifact: pr_456\n```\n\nThe runtime can preserve that relationship immediately.\n\nWhat it cannot preserve yet is an event that has not happened.\n\nSeventeen minutes later, the pull request is merged in GitHub.\n\nConceptually, the history now looks more like this:\n\n```\nAI RUNTIME\n\nexecution_123\n      ↓\nartifact: pr_456\n      │\n      │   time passes\n      │\n      ▼\n\nDOMAIN SYSTEM\n\npull_request merged\n      ↓\nartifact: pr_456\n      ↓\noutcome evidence\n```\n\nThe important part isn't GitHub specifically.\n\nThe same pattern can appear when evidence comes from a CRM update, a ticketing system, a payment event, a human review process, another workflow or any domain system that observes something the original runtime cannot know on its own.\n\nThe execution produces something.\n\nLater, another system tells us what happened to it.\n\nThat changes the engineering question.\n\nWe are no longer asking only:\n\nDid the workflow complete?\n\nWe also need to ask:\n\nCan the evidence that arrives later still be related to the execution or result it describes?\n\nSuppose the agent creates `pr_456`.\n\nIf a later event tells us that `pr_456` was merged, the relationship is relatively easy to imagine:\n\n```\nexecution_123\n      ↓\nartifact: pr_456\n      ↓\nlater domain event\n      ↓\nevidence about the outcome\n```\n\nThe identifiers here are illustrative. A real system may not need an `execution_id`, an `artifact_id` and a separate outcome identifier for every workflow.\n\nWhat matters is preserving enough relationship information that evidence arriving later does not become detached from the work it is supposed to describe.\n\nWithout that relationship, we can end up with two individually valid facts:\n\n```\nexecution_123 = COMPLETED\n```\n\nand:\n\n```\npr_456 = MERGED\n```\n\nwhile still lacking a reliable way to establish that the second fact is relevant to the first.\n\nThis is where correlation becomes more than an observability convenience.\n\nIf later domain evidence is going to influence how we interpret an execution, the system needs some defensible path between them.\n\nThat path might be direct:\n\n```\nexecution_123\n      ↓\npr_456\n```\n\nOr the domain may require more structure:\n\n```\ncustomer request\n      ↓\nlogical outcome\n      ↓\nexecution\n      ↓\ngenerated artifact\n      ↓\nlater domain evidence\n```\n\nThere is no universal shape here.\n\nIn some products, the artifact itself may provide all the identity needed. In others, several executions may contribute to one logical outcome, or one execution may produce several artifacts whose fate evolves independently.\n\nThat's why I wouldn't start by adding an `outcome_id` to every table.\n\nThe modeling question comes first:\n\n**When does the thing we're trying to evaluate become meaningfully distinct from the execution that attempted to produce it?**\n\nIf the answer is \"never,\" a separate outcome identity may add complexity without much value.\n\nBut if the execution can finish while the thing we care about continues to evolve somewhere else, treating those two concepts as independently identifiable can become useful.\n\nEven when the relationship is preserved, the evidence may not become visible to our system at the moment the underlying domain event occurs.\n\nImagine the pull request is merged at 10:17.\n\nOur system receives evidence of that merge at 10:19.\n\n```\n10:00   execution completed\n  │\n10:17   PR merged\n  │\n10:19   merge observed by our system\n```\n\nThose timestamps describe different facts.\n\nThe pull request did not become merged at 10:19 simply because that's when our system learned about it.\n\nFor the kind of reasoning we're doing here, it can be useful to preserve two temporal concepts:\n\n```\noccurred_at = 10:17\nobserved_at = 10:19\n```\n\n`occurred_at` represents when the relevant domain occurrence happened, when that information is available.\n\n`observed_at` represents when our system obtained the evidence.\n\nThe exact fields and semantics depend on the event source and architecture. The important distinction is between when something happened in the domain and when our system became able to reason from it.\n\nAt 10:18, for example, the pull request may already have been merged while our system still legitimately represents the outcome as unknown.\n\n```\nDOMAIN\nPR merged\n\nSYSTEM KNOWLEDGE\noutcome = UNKNOWN\n```\n\nAgain, there is no contradiction.\n\nThe system can only make claims from the evidence available to it.\n\nThis also means that an outcome state is not necessarily a timeless description of reality. It can represent what the system is currently justified in claiming from the evidence it has observed.\n\nAnd that becomes important as soon as the evidence changes.\n\nReturn to `pr_456`.\n\nAt 10:17, the pull request is merged.\n\nIf merge is the domain condition we're currently using to evaluate this outcome, the system may now have enough evidence to move beyond `UNKNOWN`.\n\n```\n10:00   execution completed\n  │\n10:17   PR merged\n  │\n  ▼\noutcome = CONFIRMED\n```\n\nFor the question we're asking, that may be a perfectly defensible conclusion at 10:17.\n\nNow make the example slightly harder.\n\nAt 14:42, the change is reverted.\n\n```\n10:00   execution completed\n  │\n10:17   PR merged\n  │\n14:42   change reverted\n```\n\nWhat should happen to the outcome now?\n\nThe interesting part is that the merge did not stop being a historical fact. The pull request really was merged at 10:17.\n\nWhat changed is the evidence available for whatever broader claim we're trying to make about the outcome.\n\nIf our success condition was simply:\n\nWas the pull request merged?\n\nthen the merge event may still be sufficient.\n\nIf our success condition was closer to:\n\nWas the generated change accepted and retained?\n\nthen the later revert may materially change the answer.\n\nSame execution history.\n\nNew domain evidence.\n\nPotentially different outcome interpretation.\n\nThis is another reason I'm hesitant to think of outcome success as a boolean that becomes permanently true the moment the first positive signal arrives.\n\n```\nsuccess = true\n```\n\nlooks simple.\n\nBut what does the system do when relevant evidence arrives later?\n\n```\nsuccess = false\n```\n\nOverwriting the value may tell us the latest interpretation, but it can also erase part of the history that explains how we got there.\n\nA different way to reason about the problem is to separate the evidence from the state we derive from it.\n\nConceptually:\n\n```\nOUTCOME EVIDENCE\n\n10:17   pull_request_merged\n14:42   change_reverted\n```\n\nThe system can then interpret those facts according to the success condition relevant to the product.\n\nThat does not require turning the architecture into a full event-sourced system.\n\nThe narrower point is that when evidence can arrive late or change the interpretation of an outcome, preserving the observations can be more useful than preserving only the latest boolean derived from them.\n\nInstead of remembering only:\n\n```\noutcome.success = false\n```\n\nwe may want to retain enough history to answer:\n\nWhy does the system currently consider this outcome unsuccessful?\n\nWhat evidence was available when it previously considered the outcome successful?\n\nThose are different questions from asking for the latest state.\n\nThis distinction becomes especially important when downstream systems start consuming the result.\n\nAn analytics job may have counted the outcome yesterday. A profitability calculation may have included it. An experiment may have classified it as successful.\n\nLater evidence does not necessarily mean those systems made an irrational decision at the time. They may have been operating on the evidence then available.\n\nThe problem is whether we can explain that decision after the evidence changes.\n\nThis suggests another distinction:\n\n```\nEVIDENCE\n   ↓\nINTERPRETATION\n   ↓\nCURRENT OUTCOME STATE\n```\n\nThe evidence records what was observed.\n\nThe interpretation applies the relevant success condition to that evidence.\n\nThe state represents what the system is currently justified in claiming.\n\nThose layers do not need to become three separate database tables. They are different concepts before they are implementation choices.\n\nThat matters because otherwise a field such as:\n\n```\noutcome.status = CONFIRMED\n```\n\ncan quietly become both the conclusion and the only surviving record of why that conclusion existed.\n\nOnce those ideas are separated, a reversal becomes easier to reason about.\n\nWe don't need to pretend the earlier evidence never existed. We can say that additional evidence changed the interpretation.\n\n```\n10:17\nevidence: PR merged\n     ↓\noutcome: CONFIRMED\n\n     ...\n\n14:42\nevidence: change reverted\n     ↓\noutcome: RE-EVALUATED\n```\n\n`RE-EVALUATED` is illustrative here, not a state I would prescribe for every system.\n\nThe exact state model depends on what the product means by success and which domain events are relevant to that definition.\n\nThe architectural principle is smaller:\n\n**A completed execution can have a closed execution history while the evidence used to interpret its outcome continues to evolve.**\n\nAnd that creates a consequence outside the workflow engine itself.\n\nIf outcome evidence can arrive late, or change after we have already classified an outcome, then any economic analysis built on that classification inherits the same uncertainty.\n\nSo far, none of this requires us to talk about money.\n\nThe same architecture matters for product analytics, experimentation, quality measurement and any other system that needs to reason about what happened after an AI execution.\n\nBut it becomes especially interesting when outcome state enters an economic calculation.\n\nImagine a product records 100 completed AI executions during a period.\n\nIf the analytics model silently assumes:\n\n```\nCOMPLETED = SUCCESSFUL OUTCOME\n```\n\nthen the denominator appears straightforward:\n\n```\ncompleted executions: 100\nsuccessful outcomes: 100\n```\n\nNow imagine that the relevant domain evidence arrives later.\n\nOf those 100 completed executions, 73 eventually satisfy the success condition the product is actually interested in.\n\n```\ncompleted executions:          100\nconfirmed relevant outcomes:    73\n```\n\nThese numbers are purely illustrative.\n\nThe important part is not the difference between 100 and 73. It's what the difference reveals.\n\nThe execution system may have been completely correct when it reported 100 completed executions.\n\nThe economic analysis can also be mathematically correct given the data it receives.\n\nBut if the analysis treats execution completion as evidence for a different outcome claim, the precision of the arithmetic can hide ambiguity in the denominator.\n\n**An outcome-level economic metric can be mathematically correct while still resting on a success definition the available evidence doesn't support.**\n\nThis is why I think the engineering problem has to come before the economic formula.\n\nBefore asking how much successful outcomes cost, we need to understand what the system is justified in counting as a successful outcome in the first place.\n\nThat doesn't mean execution completion is a bad metric. It answers a different question.\n\n```\nCOMPLETED EXECUTIONS\nHow many executions reached the expected terminal state?\n\nCONFIRMED OUTCOMES\nHow many outcomes currently satisfy the relevant\nsuccess condition based on available evidence?\n```\n\nDepending on the product, both may be useful.\n\nWhat becomes dangerous is allowing one to silently stand in for the other.\n\nDelayed evidence introduces another complication.\n\nAt 10:00, an outcome may still be unknown. At 10:17, new evidence may justify classifying it as confirmed. At 14:42, later evidence may force us to reconsider that classification.\n\nSo the number of outcomes considered successful is not necessarily a static property available at execution time.\n\nIt may depend on when the question is asked and which evidence was available then.\n\n```\n10:00\nexecution = COMPLETED\noutcome   = UNKNOWN\n\n    ↓\n\n10:17\nevidence  = PR merged\noutcome   = CONFIRMED\n\n    ↓\n\n14:42\nevidence  = change reverted\noutcome   = RE-EVALUATED\n```\n\nIf an economic system uses outcome state, this temporal dimension becomes part of the meaning of the metric.\n\nA historical analysis may need to distinguish between:\n\nWhat did the system have enough evidence to conclude about this outcome at the time?\n\nWhat does the evidence available now allow us to conclude about that outcome?\n\nThose questions are not necessarily interchangeable.\n\nThis is one reason I've become interested in the evidence underneath economic metrics rather than only the final numbers they produce.\n\nA number can be reproducible while the assumptions underneath its denominator remain invisible.\n\nAnd once those assumptions matter, economic observability starts looking less like a calculation problem and more like an evidence problem.\n\nIf outcome evidence can arrive after execution completes, then the runtime cannot preserve only what it knows at completion time and expect that to answer every later question.\n\nSomething has to survive long enough for the later evidence to still have meaning.\n\nNot necessarily a large schema.\n\nNot necessarily a universal outcome model.\n\nBut enough information to preserve the relationship between the work that happened and the evidence that may arrive later.\n\nFor the coding-agent example, that might mean being able to reconstruct a path like this:\n\n```\nexecution_123\n      ↓\nproduced\n      ↓\nartifact: pr_456\n      ↓\nlater evidence\n      ↓\nPR merged\n      ↓\nlater evidence\n      ↓\nchange reverted\n```\n\nThe exact identifiers are domain-specific.\n\nThe architectural requirement is more general.\n\nIf `pr_456` becomes relevant seventeen minutes after `execution_123` ends, the relationship between them cannot depend entirely on execution-local state that disappears when the worker returns.\n\nAnd if later evidence changes how the outcome is interpreted, preserving only the latest interpretation may not be enough either.\n\nAt minimum, I would want to reason about four different things:\n\n```\nEXECUTION\nWhat work happened?\n\nRESULT / ARTIFACT\nWhat did that work produce?\n\nOUTCOME EVIDENCE\nWhat did we later observe about that result?\n\nPROVENANCE\nWhere did that evidence come from, and when did we observe it?\n```\n\nThat is a conceptual model, not a proposed database schema.\n\nSome systems may collapse several of these concepts together. Others may need additional identity or lineage because multiple executions contribute to the same outcome.\n\nThe important part is that the architecture should not destroy distinctions the business may need later.\n\nThis is the part of the problem I find most interesting.\n\nThe workflow engine can legitimately say:\n\n```\nexecution.status = COMPLETED\n```\n\nand be completely correct.\n\nIt doesn't need to wait seventeen minutes for a pull request to merge before it is allowed to finish the execution. Doing so would confuse the lifecycle of the worker with the lifecycle of the domain.\n\nInstead, the execution can end while the outcome remains unresolved:\n\n```\nexecution.status = COMPLETED\noutcome.status   = UNKNOWN\n```\n\nLater evidence can move our interpretation forward without reopening the execution itself.\n\n```\nexecution.status = COMPLETED\noutcome.status   = CONFIRMED\n```\n\nAnd still later evidence may require another interpretation.\n\nThe execution history remains stable.\n\nThe evidence around the outcome continues to evolve.\n\nThat separation matters because AI workflows can interact with systems whose relevant consequences happen outside the runtime itself.\n\nA model returns. A tool finishes. An agent stops.\n\nBut the artifact it produced may continue through approval, acceptance, deployment, resolution or some other domain process.\n\nThe runtime knows when its own work ended.\n\nIt may not know what happened next.\n\nI started with a question that looked like a state-machine problem:\n\nWhat should `SUCCESS` mean when an AI workflow finishes?\n\nI now think the more interesting engineering problem begins when the answer cannot exist yet.\n\nThe execution may finish at 10:00. The evidence we care about may arrive at 10:17. Additional evidence may change the interpretation again hours later.\n\nThat creates a different boundary:\n\n```\nEXECUTION ENDS\n      │\n      │   domain continues\n      │\n      ▼\nOUTCOME EVIDENCE ARRIVES\n```\n\nOnce those timelines separate, several things follow.\n\n`UNKNOWN` becomes meaningfully different from `FAILED`. Correlation matters because later evidence needs a path back to the work it describes. Occurrence time and observation time can answer different questions. Preserving evidence may matter more than repeatedly overwriting a final boolean.\n\nI've been exploring this boundary while thinking about what economic evidence an AI monetization runtime would need when the domain answer arrives after the execution itself.\n\nIt also extends a question I explored in the latest Licenzy Guide: what evidence actually justifies calling an AI outcome successful?\n\n[Your AI Workflow Completed. Did It Actually Succeed?](https://licenzy.app/guides/ai-workflow-completed-did-it-actually-succeed)\n\nThe Guide focuses on the meaning and evidence of success.\n\nThe engineering question I've explored here starts one step later:\n\n**What happens when that evidence isn't available yet?**\n\nIf outcome state eventually becomes an input to profitability, experimentation or other economic analysis, the system has to decide when that state is stable enough to use.\n\nAnd I'm not convinced execution completion can answer that question for us.\n\nSo the question I'm left with is:\n\n**If outcome evidence can arrive late — and later change — when is an economic system justified in treating that outcome as final?**\n\n`merged == true` can be used to distinguish a merge from another close event.`time` attribute as the timestamp of when the occurrence happened.", "url": "https://wpnews.pro/news/your-ai-execution-finished-the-outcome-hasn-t-happened-yet", "canonical_source": "https://dev.to/thelastciroandrea/your-ai-execution-finished-the-outcome-hasnt-happened-yet-3885", "published_at": "2026-09-30 13:33:49+00:00", "updated_at": "2026-09-30 13:48:05.145191+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "mlops", "developer-tools"], "entities": ["GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-ai-execution-finished-the-outcome-hasn-t-happened-yet", "markdown": "https://wpnews.pro/news/your-ai-execution-finished-the-outcome-hasn-t-happened-yet.md", "text": "https://wpnews.pro/news/your-ai-execution-finished-the-outcome-hasn-t-happened-yet.txt", "jsonld": "https://wpnews.pro/news/your-ai-execution-finished-the-outcome-hasn-t-happened-yet.jsonld"}}