cd /news/ai-agents/stop-trusting-your-agent-s-final-ans… Β· home β€Ί topics β€Ί ai-agents β€Ί article
[ARTICLE Β· art-144921] src=dev.to β†— pub= topic=ai-agents verified=true sentiment=Β· neutral

Stop Trusting Your Agent's Final Answer: Build a Tiny Agent Tracer in TypeScript

A developer published a TypeScript tutorial and open-source project, tiny-agent-tracer, that records an agent run as a tree of OpenTelemetry-style spans so the execution trace can be compared against the agent's final answer. The tracer uses mocked models, tools and a virtual clock to produce deterministic output with no API key or OpenTelemetry SDK, and demonstrates a mocked agent claiming it sent an email that its own trace shows timed out.

by read11 min views4 publishedOct 4, 2026

On September 20, an OpenAI research agent was asked to identify the author of a blog post.

It ended the run politely.

"I couldn't reliably establish" who it was, it told the user (in OpenAI's English translation). It asked for the original wording, or the title.

Humble. Helpful. Nothing to see.

The trace told a different story.

OpenAI has d training, evaluation and inference with tool use for its most capable models while it closes the gap. In the post announcing the reports site, Sam Altman said the company is working to understand "petabytes of agent activity logs."

So here is my contrarian take.

The final answer is the least trustworthy thing an agent produces.

Not because agents lie.

Because the answer is a summary, written by the same system you're trying to check.

The answer is a claim. The trace is the evidence.

The tooling world clearly agrees. Look at the last two weeks:

gen_ai.skill.* attributes for the execute tool span (#498) and a gen_ai.main_agent entity (#270). The span names are already there: invoke_agent {gen_ai.agent.name}, execute_tool {gen_ai.tool.name}, {gen_ai.operation.name} {gen_ai.request.model} for a model call, and a plan span since May. All of it is still marked Development. Everyone is shipping tools to read traces.

Which only helps if your agent writes a good one.

In Evals for Agents: Did It Stay in Scope? I graded runs after the fact. This post is about recording the run well enough that there's something to grade.

Let's build a tiny tracer.

By the end, you'll run one command:

npx tsx tracer.ts

And watch a mocked agent claim it sent an email that, according to its own trace, timed out.

Smaller stakes than a DNS side channel. Same shape: the answer and the trace disagree.

No API key.

No OpenTelemetry SDK.

Just TypeScript, and the same span names the conventions use.

One honesty note: this is not how OpenAI, AWS or anyone else records traces. It's my small model of the shape their docs describe.

Code: github.com/bobbyhalljr/tiny-agent-tracer

One run.

One root span for the agent.

Children for planning, model calls, tool calls and a subagent.

Every span gets a name, a parent, a start, an end, a status and a few gen_ai.* attributes.

Then we ask the trace four questions:

It's also a small version of something I care about in Roster: an AI employee should leave evidence, not just a confident summary.

The model, the tools and every duration are a mock on a virtual clock, so the output is identical on every run.

You will need Node.js 18 or newer.

mkdir tiny-agent-tracer
cd tiny-agent-tracer

npm init -y
npm install --save-dev typescript tsx @types/node

Save the following blocks, in order, as tracer.ts.

// tiny-agent-tracer: trace one agent run with OpenTelemetry GenAI span names.
// The model, the tools and every duration are a MOCK on a virtual clock, so
// the output is the same on every run. No network. No API key.

type Attr = string | number | boolean | null;

type Span = {
  spanId: string;
  parentId: string | null;
  name: string;
  attrs: Record<string, Attr>;
  start: number;
  end: number;
  status: "ok" | "error";
};

A span is a unit of work with a parent.

That's the whole trick.

The parent link turns a flat log into a tree. The tree is what tells you that a tool call belonged to a subagent and not to the root.

Attr allows null on purpose. We'll need it for tokens.

class Tracer {
  spans: Span[] = [];
  private stack: Span[] = [];
  private now = 0;
  private nextId = 1;

  advance(ms: number) {
    this.now += ms;
  }

  span<T>(name: string, attrs: Record<string, Attr>, fn: (s: Span) => T): T {
    const parent = this.stack[this.stack.length - 1];
    const s: Span = {
      spanId: `s${this.nextId++}`,
      parentId: parent ? parent.spanId : null,
      name,
      attrs,
      start: this.now,
      end: this.now,
      status: "ok",
    };
    this.spans.push(s);
    this.stack.push(s);
    try {
      return fn(s);
    } catch (err) {
      s.status = "error";
      s.attrs["error.type"] = (err as Error).message;
      throw err;
    } finally {
      s.end = this.now;
      this.stack.pop();
    }
  }
}

span() pushes a span on a stack, runs your function, and pops it.

Whatever runs inside becomes a child.

If the function throws, the span is marked error, gets an error.type, and the error keeps going. The tracer records failures. It doesn't hide them.

The clock is fake. advance(ms) moves time forward. Real tracers read a real clock, but a fake one makes the output deterministic, which is nice for a tutorial and essential for a test.

A trace you can't reproduce is a screenshot.

const tracer = new Tracer();

function chat(model: string, ms: number, input: number | null, output: number | null) {
  return tracer.span(
    `chat ${model}`,
    {
      "gen_ai.operation.name": "chat",
      "gen_ai.request.model": model,
      "gen_ai.usage.input_tokens": input,
      "gen_ai.usage.output_tokens": output,
    },
    () => tracer.advance(ms),
  );
}

function tool(name: string, ms: number, fail?: string) {
  return tracer.span(
    `execute_tool ${name}`,
    { "gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": name },
    () => {
      tracer.advance(ms);
      if (fail) throw new Error(fail);
    },
  );
}

function agent<T>(name: string, fn: () => T): T {
  return tracer.span(
    `invoke_agent ${name}`,
    { "gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": name },
    fn,
  );
}

// MOCK run: a release-notes agent with one subagent.
const finalAnswer = agent("release-notes", () => {
  tracer.span(
    "plan release-notes",
    { "gen_ai.operation.name": "plan", "gen_ai.agent.name": "release-notes" },
    () => chat("mock-large", 820, 900, 120),
  );
  tool("read_file", 30);
  chat("mock-large", 640, 1400, 210);
  agent("changelog-checker", () => {
    chat("mock-small", 210, 600, 40);
    tool("git_log", 120);
    chat("mock-small", 180, null, null); // usage not reported yet
  });
  try {
    tool("send_email", 5000, "timeout");
  } catch {
    // the loop swallows the error and keeps going
  }
  chat("mock-large", 450, 1800, 90);
  return "Release notes drafted and sent to the team.";
});

Three helpers. Three span types.

chat records the model and token usage under gen_ai.usage.input_tokens and gen_ai.usage.output_tokens.

tool records gen_ai.tool.name.

agent records gen_ai.agent.name and wraps everything the agent does.

The mocked run plans, reads a file, calls a subagent, tries to send an email, and answers.

Two details are deliberate.

The second mock-small call reports null usage. OpenAI's guide says usage can arrive after the turn ends, and that a blank or null value "means the count is unknown. It does not mean the agent used zero tokens."

And the loop swallows the send_email timeout and keeps going. Then the final answer says "sent." That's the bug we're here to catch.

function children(id: string | null): Span[] {
  return tracer.spans.filter((s) => s.parentId === id);
}

function tokens(s: Span): string {
  const i = s.attrs["gen_ai.usage.input_tokens"];
  const o = s.attrs["gen_ai.usage.output_tokens"];
  if (s.attrs["gen_ai.operation.name"] !== "chat") return "";
  return i === null || o === null ? "  tokens=unknown" : `  in=${i} out=${o}`;
}

function printTree(id: string | null, prefix: string) {
  const kids = children(id);
  kids.forEach((s, i) => {
    const last = i === kids.length - 1;
    const branch = id === null ? "" : prefix + (last ? "└─ " : "β”œβ”€ ");
    const ms = `${s.end - s.start}ms`.padStart(7);
    const label = (branch + s.name).padEnd(40);
    const err = s.status === "error" ? `  error.type=${s.attrs["error.type"]}` : "";
    console.log(`${label} ${s.status.padEnd(5)} ${ms}${tokens(s)}${err}`);
    printTree(s.spanId, id === null ? "" : prefix + (last ? "   " : "β”‚  "));
  });
}

Depth first. Children under parents. Duration, status and tokens on every line.

tokens() prints unknown instead of 0.

Zero is a number.

Unknown is a different fact.

function ancestors(s: Span): string[] {
  const out: string[] = [];
  let p = tracer.spans.find((x) => x.spanId === s.parentId);
  while (p) {
    out.push(p.name);
    p = tracer.spans.find((x) => x.spanId === p!.parentId);
  }
  return out;
}

function failures() {
  for (const s of tracer.spans.filter((x) => x.status === "error")) {
    console.log(`  ${s.name} (${s.attrs["error.type"]}) <- ${ancestors(s).join(" <- ")}`);
  }
}

function claimCheck(answer: string) {
  const emailFailed = tracer.spans.some(
    (s) => s.attrs["gen_ai.tool.name"] === "send_email" && s.status === "error",
  );
  if (answer.includes("sent") && emailFailed) {
    console.log(`  answer says "sent", but execute_tool send_email failed`);
  }
}

function timeByOperation() {
  const total = new Map<string, number>();
  for (const s of tracer.spans) {
    const op = String(s.attrs["gen_ai.operation.name"]);
    if (op === "invoke_agent" || op === "plan") continue; // parents, not work
    total.set(op, (total.get(op) ?? 0) + (s.end - s.start));
  }
  for (const [op, ms] of total) console.log(`  ${op.padEnd(13)} ${ms}ms`);
}

function tokensByAgent() {
  for (const a of tracer.spans.filter((s) => s.attrs["gen_ai.operation.name"] === "invoke_agent")) {
    let known = 0;
    let unknown = 0;
    const own = (id: string): Span[] =>
      children(id).flatMap((c) =>
        c.attrs["gen_ai.operation.name"] === "invoke_agent" ? [] : [c, ...own(c.spanId)],
      );
    for (const c of own(a.spanId)) {
      if (c.attrs["gen_ai.operation.name"] !== "chat") continue;
      const i = c.attrs["gen_ai.usage.input_tokens"];
      const o = c.attrs["gen_ai.usage.output_tokens"];
      if (i === null || o === null) unknown++;
      else known += Number(i) + Number(o);
    }
    const note = unknown ? `, ${unknown} chat span unknown` : "";
    console.log(`  ${String(a.attrs["gen_ai.agent.name"]).padEnd(18)} ${known} tokens${note}`);
  }
}

failures() walks up from every failed span, so you see the error and everything it happened under.

claimCheck() is crude on purpose. If the answer says "sent" and the email tool failed, it says so. A real version would compare structured claims against tool results. The idea is the same.

timeByOperation() sums durations by gen_ai.operation.name. It skips invoke_agent and plan, because those are parents. Counting them would count the same time twice.

tokensByAgent() counts each agent's own model calls and stops at subagent boundaries. That matches OpenAI's guide: an agent span's usage covers the agent itself, "they do not include its subagents."

console.log("Trace (MOCK model, virtual clock)\n");
printTree(null, "");
console.log(`\nFinal answer: "${finalAnswer}"`);
console.log("\nWhat failed, and under what?");
failures();
console.log("\nDoes the answer match the trace?");
claimCheck(finalAnswer);
console.log("\nWhere did the time go?");
timeByOperation();
console.log("\nTokens per agent (subagents counted separately):");
tokensByAgent();

Run it:

npx tsx tracer.ts

You should see:

Trace (MOCK model, virtual clock)

invoke_agent release-notes               ok     7450ms
β”œβ”€ plan release-notes                    ok      820ms
β”‚  └─ chat mock-large                    ok      820ms  in=900 out=120
β”œβ”€ execute_tool read_file                ok       30ms
β”œβ”€ chat mock-large                       ok      640ms  in=1400 out=210
β”œβ”€ invoke_agent changelog-checker        ok      510ms
β”‚  β”œβ”€ chat mock-small                    ok      210ms  in=600 out=40
β”‚  β”œβ”€ execute_tool git_log               ok      120ms
β”‚  └─ chat mock-small                    ok      180ms  tokens=unknown
β”œβ”€ execute_tool send_email               error  5000ms  error.type=timeout
└─ chat mock-large                       ok      450ms  in=1800 out=90

Final answer: "Release notes drafted and sent to the team."

What failed, and under what?
  execute_tool send_email (timeout) <- invoke_agent release-notes

Does the answer match the trace?
  answer says "sent", but execute_tool send_email failed

Where did the time go?
  chat          2300ms
  execute_tool  5150ms

Tokens per agent (subagents counted separately):
  release-notes      4520 tokens
  changelog-checker  640 tokens, 1 chat span unknown

The final answer is confident.

The trace disagrees.

send_email ran for 5000ms and ended in error.type=timeout. The root agent span still says ok, because the loop caught the error. A run-level status would have told you nothing.

That's the lesson from OpenAI's monitor, in miniature. A failed attempt is still an attempt. The span records that it happened, not just whether it worked.

The time question has a boring answer, which is the best kind. 5150 of 7450ms went to tool calls, and 5000 of those were one email that never went out.

And the subagent's tokens come with a footnote. 640 known, 1 call unknown. Not 640. Not zero. At least 640.

The answer is what the agent says. The trace is what the agent did.

This is a teaching tracer. Here is what a real one would need.

We keep spans in an array. Real systems use an OpenTelemetry SDK with context propagation, sampling and an OTLP exporter. The span names and attributes are the part worth copying.

A stack works for synchronous code. Two subagents running at once need real context propagation, or every span ends up under whichever one started last.

We don't record prompts or tool arguments. The conventions treat message content as opt-in for a reason. Traces of agents that read inboxes are full of other people's data.

OpenAI's report says DNS activity was logged, but an infrastructure detector for anomalous DNS excluded the affected environment. A trace nobody queries is a very detailed diary. Our four questions are hard-coded. Real ones belong in alerts that can stop a run.

Everything gen_ai.* here is Development status. Two changes landed this week alone. Pin a version and expect renames.

Logs answer: what happened?

Traces answer: what happened, under what, for how long, and at whose cost?

For a chatbot, the transcript was enough.

For an agent with tools, subagents and a budget, it isn't.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ invoke_agent                             β”‚
β”‚   β”œβ”€β”€ plan ──→ chat                      β”‚
β”‚   β”œβ”€β”€ execute_tool                       β”‚
β”‚   β”œβ”€β”€ invoke_agent (subagent)            β”‚
β”‚   β”‚     β”œβ”€β”€ chat                         β”‚
β”‚   β”‚     └── execute_tool                 β”‚
β”‚   └── chat ──→ final answer (a claim)    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                     ↓
        failures Β· claims Β· time Β· tokens

The span provides a unit of work.

The parent link provides structure.

The status provides the truth about each step.

The attributes provide a shared vocabulary.

The clock provides cost in time.

The usage provides cost in tokens, or an honest "unknown."

The final answer provides a claim to check.

If your agent can't show its work, you're grading its confidence.

I'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access. They work inside a lane, leave a record of what they did, and ask before doing anything you'd want to see first.

If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.

── more in #ai-agents 4 stories Β· sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/stop-trusting-your-a…] indexed:0 read:11min 2026-10-04 Β· β€”