cd /news/ai-agents/your-ai-execution-finished-the-outco… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-142529] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

Your AI Execution Finished. The Outcome Hasn't Happened Yet

A developer argues that AI coding agents' terminal execution states like COMPLETED or SUCCESS describe only the runtime's own lifecycle, not the eventual real-world outcome, since evidence such as a pull request being merged can arrive minutes later from another system. The writeup proposes separating the execution timeline from the outcome-evidence timeline, so that execution.status = COMPLETED can coexist with outcome.status = UNKNOWN without implying failure.

by read16 min views2 publishedSep 30, 2026

An AI coding agent finishes its work at 10:00.

It generated the requested change, returned the result and reached the terminal state we expected.

10:00

execution_123
  status: COMPLETED

From the execution's point of view, there may be nothing left to do.

But imagine that the result is a pull request.

At 10:02, the automated tests finish successfully. At 10:08, a human reviewer approves the change. At 10:17, the pull request is finally merged.

10:00  agent completed
   β”‚
10:02  tests passed
   β”‚
10:08  human approved
   β”‚
10:17  PR merged

Now look again at what we knew at 10:00.

The execution had finished. The event that would later tell us the pull request was actually merged had not happened yet.

That creates an architectural problem I find more interesting than simply asking whether COMPLETED means SUCCESS.

What happens when the evidence we need to evaluate an outcome belongs to another system β€” and arrives after the execution itself is already over?

I've been exploring this question while thinking about the economic evidence needed underneath AI monetization infrastructure.

The useful distinction, at least for me, is not that workflow completion is somehow a weak or incorrect state. COMPLETED may describe the execution perfectly.

The problem appears when we expect that execution state to also tell us something that can only become knowable later.

A workflow state machine often has a shape like this:

PENDING
   ↓
RUNNING
   ↓
COMPLETED

Or perhaps:

      β”Œβ”€β”€ FAILED
      β”‚
RUNNING
      β”‚
      └── COMPLETED

Those states answer an important question:

What happened to the execution?

Did it start? Is it still running? Did it terminate? Did the runtime encounter a failure?

Nothing about that model is inherently incomplete.

The problem begins when another question gets attached to the same terminal state:

Did the intended outcome happen?

For our coding agent, those questions can exist on different timelines.

EXECUTION

10:00
REQUEST
   ↓
RUNNING
   ↓
COMPLETED
   β”‚
   β”‚   outcome not yet known
   β”‚
   β–Ό
10:17
PR MERGED

At 10:00, the runtime cannot observe an event that will not occur until 10:17.

That sounds obvious when written as a timeline. It becomes less obvious when the data model compresses both ideas into something like:

status = SUCCESS

What does SUCCESS mean here?

Did the agent successfully finish? Did the generated change pass validation? Was the artifact accepted? Was the pull request merged?

These can all be legitimate success conditions. They just don't describe the same fact.

A product may intentionally define execution success as successful technical completion. There is nothing wrong with that.

The ambiguity appears when the same state is later reused as evidence that a different domain condition was satisfied.

This is the mental model I've found more useful:

EXECUTION TIMELINE
        β‰ 
OUTCOME-EVIDENCE TIMELINE

I'm using those terms descriptively rather than as a formal distributed-systems taxonomy.

The point is simply that execution state and evidence about the resulting outcome may evolve independently.

That means this state is not contradictory:

execution.status = COMPLETED
outcome.status   = UNKNOWN

UNKNOWN does not mean the outcome failed.

It means the system does not yet have enough evidence to make the outcome claim we're asking it to make.

Not yet known is not the same as failed.

And once that evidence can arrive later, the problem stops being only about workflow state.

It becomes a distributed-systems problem.

Go back to the coding agent.

At 10:00, its execution completes and produces a pull request.

execution_123
      ↓
artifact: pr_456

The runtime can preserve that relationship immediately.

What it cannot preserve yet is an event that has not happened.

Seventeen minutes later, the pull request is merged in GitHub.

Conceptually, the history now looks more like this:

AI RUNTIME

execution_123
      ↓
artifact: pr_456
      β”‚
      β”‚   time passes
      β”‚
      β–Ό

DOMAIN SYSTEM

pull_request merged
      ↓
artifact: pr_456
      ↓
outcome evidence

The important part isn't GitHub specifically.

The same pattern can appear when evidence comes from a CRM update, a ticketing system, a payment event, a human review process, another workflow or any domain system that observes something the original runtime cannot know on its own.

The execution produces something.

Later, another system tells us what happened to it.

That changes the engineering question.

We are no longer asking only:

Did the workflow complete?

We also need to ask:

Can the evidence that arrives later still be related to the execution or result it describes?

Suppose the agent creates pr_456.

If a later event tells us that pr_456 was merged, the relationship is relatively easy to imagine:

execution_123
      ↓
artifact: pr_456
      ↓
later domain event
      ↓
evidence about the outcome

The identifiers here are illustrative. A real system may not need an execution_id, an artifact_id and a separate outcome identifier for every workflow.

What matters is preserving enough relationship information that evidence arriving later does not become detached from the work it is supposed to describe.

Without that relationship, we can end up with two individually valid facts:

execution_123 = COMPLETED

and:

pr_456 = MERGED

while still lacking a reliable way to establish that the second fact is relevant to the first.

This is where correlation becomes more than an observability convenience.

If later domain evidence is going to influence how we interpret an execution, the system needs some defensible path between them.

That path might be direct:

execution_123
      ↓
pr_456

Or the domain may require more structure:

customer request
      ↓
logical outcome
      ↓
execution
      ↓
generated artifact
      ↓
later domain evidence

There is no universal shape here.

In some products, the artifact itself may provide all the identity needed. In others, several executions may contribute to one logical outcome, or one execution may produce several artifacts whose fate evolves independently.

That's why I wouldn't start by adding an outcome_id to every table.

The modeling question comes first:

When does the thing we're trying to evaluate become meaningfully distinct from the execution that attempted to produce it?

If the answer is "never," a separate outcome identity may add complexity without much value.

But if the execution can finish while the thing we care about continues to evolve somewhere else, treating those two concepts as independently identifiable can become useful.

Even when the relationship is preserved, the evidence may not become visible to our system at the moment the underlying domain event occurs.

Imagine the pull request is merged at 10:17.

Our system receives evidence of that merge at 10:19.

10:00   execution completed
  β”‚
10:17   PR merged
  β”‚
10:19   merge observed by our system

Those timestamps describe different facts.

The pull request did not become merged at 10:19 simply because that's when our system learned about it.

For the kind of reasoning we're doing here, it can be useful to preserve two temporal concepts:

occurred_at = 10:17
observed_at = 10:19

occurred_at represents when the relevant domain occurrence happened, when that information is available.

observed_at represents when our system obtained the evidence.

The exact fields and semantics depend on the event source and architecture. The important distinction is between when something happened in the domain and when our system became able to reason from it.

At 10:18, for example, the pull request may already have been merged while our system still legitimately represents the outcome as unknown.

DOMAIN
PR merged

SYSTEM KNOWLEDGE
outcome = UNKNOWN

Again, there is no contradiction.

The system can only make claims from the evidence available to it.

This also means that an outcome state is not necessarily a timeless description of reality. It can represent what the system is currently justified in claiming from the evidence it has observed.

And that becomes important as soon as the evidence changes.

Return to pr_456.

At 10:17, the pull request is merged.

If merge is the domain condition we're currently using to evaluate this outcome, the system may now have enough evidence to move beyond UNKNOWN.

10:00   execution completed
  β”‚
10:17   PR merged
  β”‚
  β–Ό
outcome = CONFIRMED

For the question we're asking, that may be a perfectly defensible conclusion at 10:17.

Now make the example slightly harder.

At 14:42, the change is reverted.

10:00   execution completed
  β”‚
10:17   PR merged
  β”‚
14:42   change reverted

What should happen to the outcome now?

The interesting part is that the merge did not stop being a historical fact. The pull request really was merged at 10:17.

What changed is the evidence available for whatever broader claim we're trying to make about the outcome.

If our success condition was simply:

Was the pull request merged?

then the merge event may still be sufficient.

If our success condition was closer to:

Was the generated change accepted and retained?

then the later revert may materially change the answer.

Same execution history.

New domain evidence.

Potentially different outcome interpretation.

This is another reason I'm hesitant to think of outcome success as a boolean that becomes permanently true the moment the first positive signal arrives.

success = true

looks simple.

But what does the system do when relevant evidence arrives later?

success = false

Overwriting the value may tell us the latest interpretation, but it can also erase part of the history that explains how we got there.

A different way to reason about the problem is to separate the evidence from the state we derive from it.

Conceptually:

OUTCOME EVIDENCE

10:17   pull_request_merged
14:42   change_reverted

The system can then interpret those facts according to the success condition relevant to the product.

That does not require turning the architecture into a full event-sourced system.

The narrower point is that when evidence can arrive late or change the interpretation of an outcome, preserving the observations can be more useful than preserving only the latest boolean derived from them.

Instead of remembering only:

outcome.success = false

we may want to retain enough history to answer:

Why does the system currently consider this outcome unsuccessful?

What evidence was available when it previously considered the outcome successful?

Those are different questions from asking for the latest state.

This distinction becomes especially important when downstream systems start consuming the result.

An analytics job may have counted the outcome yesterday. A profitability calculation may have included it. An experiment may have classified it as successful.

Later evidence does not necessarily mean those systems made an irrational decision at the time. They may have been operating on the evidence then available.

The problem is whether we can explain that decision after the evidence changes.

This suggests another distinction:

EVIDENCE
   ↓
INTERPRETATION
   ↓
CURRENT OUTCOME STATE

The evidence records what was observed.

The interpretation applies the relevant success condition to that evidence.

The state represents what the system is currently justified in claiming.

Those layers do not need to become three separate database tables. They are different concepts before they are implementation choices.

That matters because otherwise a field such as:

outcome.status = CONFIRMED

can quietly become both the conclusion and the only surviving record of why that conclusion existed.

Once those ideas are separated, a reversal becomes easier to reason about.

We don't need to pretend the earlier evidence never existed. We can say that additional evidence changed the interpretation.

10:17
evidence: PR merged
     ↓
outcome: CONFIRMED

     ...

14:42
evidence: change reverted
     ↓
outcome: RE-EVALUATED

RE-EVALUATED is illustrative here, not a state I would prescribe for every system.

The exact state model depends on what the product means by success and which domain events are relevant to that definition.

The architectural principle is smaller:

A completed execution can have a closed execution history while the evidence used to interpret its outcome continues to evolve.

And that creates a consequence outside the workflow engine itself.

If outcome evidence can arrive late, or change after we have already classified an outcome, then any economic analysis built on that classification inherits the same uncertainty.

So far, none of this requires us to talk about money.

The same architecture matters for product analytics, experimentation, quality measurement and any other system that needs to reason about what happened after an AI execution.

But it becomes especially interesting when outcome state enters an economic calculation.

Imagine a product records 100 completed AI executions during a period.

If the analytics model silently assumes:

COMPLETED = SUCCESSFUL OUTCOME

then the denominator appears straightforward:

completed executions: 100
successful outcomes: 100

Now imagine that the relevant domain evidence arrives later.

Of those 100 completed executions, 73 eventually satisfy the success condition the product is actually interested in.

completed executions:          100
confirmed relevant outcomes:    73

These numbers are purely illustrative.

The important part is not the difference between 100 and 73. It's what the difference reveals.

The execution system may have been completely correct when it reported 100 completed executions.

The economic analysis can also be mathematically correct given the data it receives.

But if the analysis treats execution completion as evidence for a different outcome claim, the precision of the arithmetic can hide ambiguity in the denominator.

An outcome-level economic metric can be mathematically correct while still resting on a success definition the available evidence doesn't support.

This is why I think the engineering problem has to come before the economic formula.

Before asking how much successful outcomes cost, we need to understand what the system is justified in counting as a successful outcome in the first place.

That doesn't mean execution completion is a bad metric. It answers a different question.

COMPLETED EXECUTIONS
How many executions reached the expected terminal state?

CONFIRMED OUTCOMES
How many outcomes currently satisfy the relevant
success condition based on available evidence?

Depending on the product, both may be useful.

What becomes dangerous is allowing one to silently stand in for the other.

Delayed evidence introduces another complication.

At 10:00, an outcome may still be unknown. At 10:17, new evidence may justify classifying it as confirmed. At 14:42, later evidence may force us to reconsider that classification.

So the number of outcomes considered successful is not necessarily a static property available at execution time.

It may depend on when the question is asked and which evidence was available then.

10:00
execution = COMPLETED
outcome   = UNKNOWN

    ↓

10:17
evidence  = PR merged
outcome   = CONFIRMED

    ↓

14:42
evidence  = change reverted
outcome   = RE-EVALUATED

If an economic system uses outcome state, this temporal dimension becomes part of the meaning of the metric.

A historical analysis may need to distinguish between:

What did the system have enough evidence to conclude about this outcome at the time?

What does the evidence available now allow us to conclude about that outcome?

Those questions are not necessarily interchangeable.

This is one reason I've become interested in the evidence underneath economic metrics rather than only the final numbers they produce.

A number can be reproducible while the assumptions underneath its denominator remain invisible.

And once those assumptions matter, economic observability starts looking less like a calculation problem and more like an evidence problem.

If outcome evidence can arrive after execution completes, then the runtime cannot preserve only what it knows at completion time and expect that to answer every later question.

Something has to survive long enough for the later evidence to still have meaning.

Not necessarily a large schema.

Not necessarily a universal outcome model.

But enough information to preserve the relationship between the work that happened and the evidence that may arrive later.

For the coding-agent example, that might mean being able to reconstruct a path like this:

execution_123
      ↓
produced
      ↓
artifact: pr_456
      ↓
later evidence
      ↓
PR merged
      ↓
later evidence
      ↓
change reverted

The exact identifiers are domain-specific.

The architectural requirement is more general.

If pr_456 becomes relevant seventeen minutes after execution_123 ends, the relationship between them cannot depend entirely on execution-local state that disappears when the worker returns.

And if later evidence changes how the outcome is interpreted, preserving only the latest interpretation may not be enough either.

At minimum, I would want to reason about four different things:

EXECUTION
What work happened?

RESULT / ARTIFACT
What did that work produce?

OUTCOME EVIDENCE
What did we later observe about that result?

PROVENANCE
Where did that evidence come from, and when did we observe it?

That is a conceptual model, not a proposed database schema.

Some systems may collapse several of these concepts together. Others may need additional identity or lineage because multiple executions contribute to the same outcome.

The important part is that the architecture should not destroy distinctions the business may need later.

This is the part of the problem I find most interesting.

The workflow engine can legitimately say:

execution.status = COMPLETED

and be completely correct.

It doesn't need to wait seventeen minutes for a pull request to merge before it is allowed to finish the execution. Doing so would confuse the lifecycle of the worker with the lifecycle of the domain.

Instead, the execution can end while the outcome remains unresolved:

execution.status = COMPLETED
outcome.status   = UNKNOWN

Later evidence can move our interpretation forward without reopening the execution itself.

execution.status = COMPLETED
outcome.status   = CONFIRMED

And still later evidence may require another interpretation.

The execution history remains stable.

The evidence around the outcome continues to evolve.

That separation matters because AI workflows can interact with systems whose relevant consequences happen outside the runtime itself.

A model returns. A tool finishes. An agent stops.

But the artifact it produced may continue through approval, acceptance, deployment, resolution or some other domain process.

The runtime knows when its own work ended.

It may not know what happened next.

I started with a question that looked like a state-machine problem:

What should SUCCESS mean when an AI workflow finishes?

I now think the more interesting engineering problem begins when the answer cannot exist yet.

The execution may finish at 10:00. The evidence we care about may arrive at 10:17. Additional evidence may change the interpretation again hours later.

That creates a different boundary:

EXECUTION ENDS
      β”‚
      β”‚   domain continues
      β”‚
      β–Ό
OUTCOME EVIDENCE ARRIVES

Once those timelines separate, several things follow.

UNKNOWN becomes meaningfully different from FAILED. Correlation matters because later evidence needs a path back to the work it describes. Occurrence time and observation time can answer different questions. Preserving evidence may matter more than repeatedly overwriting a final boolean.

I've been exploring this boundary while thinking about what economic evidence an AI monetization runtime would need when the domain answer arrives after the execution itself.

It also extends a question I explored in the latest Licenzy Guide: what evidence actually justifies calling an AI outcome successful?

Your AI Workflow Completed. Did It Actually Succeed?

The Guide focuses on the meaning and evidence of success.

The engineering question I've explored here starts one step later:

What happens when that evidence isn't available yet?

If outcome state eventually becomes an input to profitability, experimentation or other economic analysis, the system has to decide when that state is stable enough to use.

And I'm not convinced execution completion can answer that question for us.

So the question I'm left with is:

If outcome evidence can arrive late β€” and later change β€” when is an economic system justified in treating that outcome as final?

merged == true can be used to distinguish a merge from another close event.time attribute as the timestamp of when the occurrence happened.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @github 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/your-ai-execution-fi…] indexed:0 read:16min 2026-09-30 Β· β€”