cd /news/developer-tools/execution-trees-not-more-logs-a-bett… · home topics developer-tools article
[ARTICLE · art-118591] src=dev.to ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Execution Trees, Not More Logs: A Better Debugging Model for AI Agents

A developer introduced execution trees as a debugging model for AI agents in the open-source toolkit AgentInspect, arguing that flat logs fail to capture causality in complex agent runs. The tool records nested step relationships, making failures and fallbacks explicit, and is available as a TypeScript library with a local inspection CLI.

read5 min views1 publishedSep 2, 2026

A flat log can tell you that five things happened. It often cannot tell you which operation caused the next one, which failure triggered a fallback, or whether three tool calls were children of one planning step or unrelated work.

That distinction matters for AI agents because the path is part of the behavior.

I maintain AgentInspect, an open-source TypeScript toolkit for inspecting agent executions locally. This article explains why I chose execution trees as the primary debugging model, using synthetic fixtures verified against agent-inspect@6.17.4.

Consider a support agent that performs these operations:

09:00:00.000 plan started
09:00:00.020 inventory request started
09:00:00.060 inventory request failed: 503
09:00:00.061 inventory request started
09:00:00.120 inventory request succeeded
09:00:00.150 answer completed

This is enough to reconstruct a simple story, but the reconstruction is happening in your head. Add nested agents, parallel tools, reused operation names, and interleaved application logs, and timestamps stop being a reliable picture of causality.

An execution tree makes the relationship explicit:

support-agent
├── plan
├── fetch-inventory (failed: 503)
├── fetch-inventory (success)
└── draft-answer

The tree does not replace raw event data. It is a projection of that data for the question developers usually ask first: What path did this run take?

AgentInspect provides wrappers for a run and for named steps. Here is a deliberately small example:

import { inspectRun, step } from "agent-inspect";

await inspectRun(
  "travel-planner",
  async () => {
    const plan = await step("plan", async () => ({
      destinations: ["SFO", "SEA"],
    }));

    const [flights, hotels] = await Promise.all([
      step.tool("search-flights", async () => [
        { id: "F-101", price: 220 },
      ]),
      step.tool("search-hotels", async () => [
        { id: "H-202", nightly: 180 },
      ]),
    ]);

    return step.llm("rank-options", async () => ({
      plan,
      flights,
      hotels,
    }));
  },
  { traceDir: "./.agent-inspect" },
);

This is manual instrumentation. It does not claim that a wrapper can automatically discover every framework-internal operation. The purpose is to record the boundaries you care about: the run, its planning step, the two sibling tool calls, and the final model-facing step.

Then inspect the run locally:

npx agent-inspect view travel-planner \
  --dir .agent-inspect \
  --summary

A three-level synthetic fixture renders like this:

Execution Tree:
✔ outer (120ms)
  ✔ middle (80ms)
    ✔ inner (50ms)

Those two spaces are not decoration. They tell us that inner

belongs to middle

, which belongs to outer

. If inner

fails, we know which higher-level operation owned it. With flat logs, matching IDs or surrounding timestamps would be required to infer the same structure.

Nesting is especially useful when one agent delegates to another, a tool performs several sub-operations, or a retrieval step owns both a query rewrite and a vector search.

Now consider an error-recovery fixture:

Execution Tree:
✖ tool:primary-search (100ms)
    Error: primary search unavailable
✔ tool:fallback-search (200ms)
✔ handle-recovered-result (50ms)

The final run may still be successful. If we looked only at the answer, the failed primary search could disappear from the debugging story. The tree preserves both facts:

That distinction can change the engineering decision. A successful answer produced by a fallback may be acceptable, but a sudden rise in fallback use could still indicate a degraded dependency or an expensive routing change.

Retries deserve their own visible shape:

Execution Tree:
✖ tool:fetch-inventory (40ms)
    Error: synthetic 503 from upstream
✖ tool:fetch-inventory (45ms)
    Error: synthetic 503 from upstream
✔ tool:fetch-inventory (60ms)
✔ handle-recovered-result (30ms)

A final success status would hide the cost of reaching success. The repeated tool name makes the retry sequence visible. It also gives a deterministic check something concrete to evaluate: for example, whether fetch-inventory

exceeded an allowed call count.

The tree alone does not tell us whether the retry policy was correct. It gives us evidence that the policy was exercised.

A parallel fixture renders as sibling operations:

Execution Tree:
✔ tool:search-hotels (300ms)
✔ tool:search-flights (200ms)
✔ tool:search-cars (100ms)

The durations are not meant to be added. These steps are siblings, and may overlap. That protects us from a common timeline mistake: assuming each timestamped operation waited for the previous one.

The tree does not prove that concurrency was optimally implemented, but it accurately preserves the structural relationship needed to investigate it.

It is tempting to turn a readable tree into the only stored artifact. I avoided that because a human-readable view necessarily compresses information.

The underlying trace may include identifiers, timestamps, status, inputs or outputs (subject to capture policy), observations, and metadata. Different questions need different projections:

structured trace
├── tree      -> what path happened?
├── check     -> did an invariant hold?
├── diff      -> what changed between runs?
├── report    -> what should a reviewer read?
└── bundle    -> what evidence can be shared?

An execution tree is the fastest entry point, not a substitute for checks or analysis.

Suppose the retry tree reveals that an inventory tool can run three times. If the intended policy permits at most two calls, encode that expectation rather than relying on future visual inspection.

At the CLI level, a trajectory check can require tools and fail on recorded observations:

npx agent-inspect check travel-planner \
  --dir .agent-inspect \
  --preset trajectory \
  --required-tool search-flights \
  --fail-on-observation failed

For richer rules, AgentInspect exposes an experimental TraceContract

API that can express tool requirements, forbidden tools, maximum calls, ordering, run status, duration, model allowlists, and token ceilings. Because that API is beta in the referenced release, pin the version and test the exact semantics before using it as a CI gate.

The important workflow is broader than one API:

A clean tree does not prove that an answer is correct. A required retrieval step may return irrelevant documents. A model call may produce unsupported claims. A tool can succeed technically while returning stale data.

Execution trees are strongest for structural questions:

Use semantic evaluators, domain tests, and human review for content quality. The most reliable agent debugging workflow combines these layers rather than asking one visualization to answer every question.

The final response is what the user sees, but the execution path is what the engineer can improve. A tree turns that path from an inferred narrative into a concrete artifact.

That is the design principle behind AgentInspect’s local view: preserve causal structure, expose unsuccessful work even when recovery succeeds, and make suspicious patterns easy to convert into repeatable checks.

You can explore the exact release used here on GitHub. If you try it, start with a synthetic failure-and-fallback fixture. A perfect happy path is the least interesting test of a debugger.

── more in #developer-tools 4 stories · sorted by recency
── more on @agentinspect 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/execution-trees-not-…] indexed:0 read:5min 2026-09-02 ·