{"slug": "stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript", "title": "Stop Trusting Your Agent's Final Answer: Build a Tiny Agent Tracer in TypeScript", "summary": "A developer published a TypeScript tutorial and open-source project, tiny-agent-tracer, that records an agent run as a tree of OpenTelemetry-style spans so the execution trace can be compared against the agent's final answer. The tracer uses mocked models, tools and a virtual clock to produce deterministic output with no API key or OpenTelemetry SDK, and demonstrates a mocked agent claiming it sent an email that its own trace shows timed out.", "body_md": "On September 20, an OpenAI research agent was asked to identify the author of a blog post.\n\nIt ended the run politely.\n\n\"I couldn't reliably establish\" who it was, it told the user (in OpenAI's English translation). It asked for the original wording, or the title.\n\nHumble. Helpful. Nothing to see.\n\nThe trace told a different story.\n\nOpenAI has paused training, evaluation and inference with tool use for its most capable models while it closes the gap. In the [post announcing the reports site](https://techcrunch.com/2026/09/28/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity/), Sam Altman said the company is working to understand \"petabytes of agent activity logs.\"\n\nSo here is my contrarian take.\n\nThe final answer is the least trustworthy thing an agent produces.\n\nNot because agents lie.\n\nBecause the answer is a summary, written by the same system you're trying to check.\n\n**The answer is a claim. The trace is the evidence.**\n\nThe tooling world clearly agrees. Look at the last two weeks:\n\n`gen_ai.skill.*` attributes for the execute tool span (#498) and a `gen_ai.main_agent` entity (#270). The span names are already there: `invoke_agent {gen_ai.agent.name}`, `execute_tool {gen_ai.tool.name}`, `{gen_ai.operation.name} {gen_ai.request.model}` for a model call, and a `plan` span since May. All of it is still marked Development.\nEveryone is shipping tools to read traces.\n\nWhich only helps if your agent writes a good one.\n\nIn [Evals for Agents: Did It Stay in Scope?](https://dev.to/bobbyhalljr/evals-for-agents-did-it-stay-in-scope-build-a-tiny-one-in-typescript-50j2) I graded runs after the fact. This post is about recording the run well enough that there's something to grade.\n\nLet's build a tiny tracer.\n\nBy the end, you'll run one command:\n\n```\nnpx tsx tracer.ts\n```\n\nAnd watch a mocked agent claim it sent an email that, according to its own trace, timed out.\n\nSmaller stakes than a DNS side channel. Same shape: the answer and the trace disagree.\n\nNo API key.\n\nNo OpenTelemetry SDK.\n\nJust TypeScript, and the same span names the conventions use.\n\nOne honesty note: this is not how OpenAI, AWS or anyone else records traces. It's my small model of the shape their docs describe.\n\n**Code:** [github.com/bobbyhalljr/tiny-agent-tracer](https://github.com/bobbyhalljr/tiny-agent-tracer)\n\nOne run.\n\nOne root span for the agent.\n\nChildren for planning, model calls, tool calls and a subagent.\n\nEvery span gets a name, a parent, a start, an end, a status and a few `gen_ai.*` attributes.\n\nThen we ask the trace four questions:\n\nIt's also a small version of something I care about in [Roster](https://get-roster.com): an AI employee should leave evidence, not just a confident summary.\n\nThe model, the tools and every duration are a mock on a virtual clock, so the output is identical on every run.\n\nYou will need Node.js 18 or newer.\n\n```\nmkdir tiny-agent-tracer\ncd tiny-agent-tracer\n\nnpm init -y\nnpm install --save-dev typescript tsx @types/node\n```\n\nSave the following blocks, in order, as `tracer.ts`.\n\n```\n// tiny-agent-tracer: trace one agent run with OpenTelemetry GenAI span names.\n// The model, the tools and every duration are a MOCK on a virtual clock, so\n// the output is the same on every run. No network. No API key.\n\ntype Attr = string | number | boolean | null;\n\ntype Span = {\n  spanId: string;\n  parentId: string | null;\n  name: string;\n  attrs: Record<string, Attr>;\n  start: number;\n  end: number;\n  status: \"ok\" | \"error\";\n};\n```\n\nA span is a unit of work with a parent.\n\nThat's the whole trick.\n\nThe parent link turns a flat log into a tree. The tree is what tells you that a tool call belonged to a subagent and not to the root.\n\n`Attr` allows `null` on purpose. We'll need it for tokens.\n\n```\nclass Tracer {\n  spans: Span[] = [];\n  private stack: Span[] = [];\n  private now = 0;\n  private nextId = 1;\n\n  advance(ms: number) {\n    this.now += ms;\n  }\n\n  span<T>(name: string, attrs: Record<string, Attr>, fn: (s: Span) => T): T {\n    const parent = this.stack[this.stack.length - 1];\n    const s: Span = {\n      spanId: `s${this.nextId++}`,\n      parentId: parent ? parent.spanId : null,\n      name,\n      attrs,\n      start: this.now,\n      end: this.now,\n      status: \"ok\",\n    };\n    this.spans.push(s);\n    this.stack.push(s);\n    try {\n      return fn(s);\n    } catch (err) {\n      s.status = \"error\";\n      s.attrs[\"error.type\"] = (err as Error).message;\n      throw err;\n    } finally {\n      s.end = this.now;\n      this.stack.pop();\n    }\n  }\n}\n```\n\n`span()` pushes a span on a stack, runs your function, and pops it.\n\nWhatever runs inside becomes a child.\n\nIf the function throws, the span is marked `error`, gets an `error.type`, and the error keeps going. The tracer records failures. It doesn't hide them.\n\nThe clock is fake. `advance(ms)` moves time forward. Real tracers read a real clock, but a fake one makes the output deterministic, which is nice for a tutorial and essential for a test.\n\n**A trace you can't reproduce is a screenshot.**\n\n``` js\nconst tracer = new Tracer();\n\nfunction chat(model: string, ms: number, input: number | null, output: number | null) {\n  return tracer.span(\n    `chat ${model}`,\n    {\n      \"gen_ai.operation.name\": \"chat\",\n      \"gen_ai.request.model\": model,\n      \"gen_ai.usage.input_tokens\": input,\n      \"gen_ai.usage.output_tokens\": output,\n    },\n    () => tracer.advance(ms),\n  );\n}\n\nfunction tool(name: string, ms: number, fail?: string) {\n  return tracer.span(\n    `execute_tool ${name}`,\n    { \"gen_ai.operation.name\": \"execute_tool\", \"gen_ai.tool.name\": name },\n    () => {\n      tracer.advance(ms);\n      if (fail) throw new Error(fail);\n    },\n  );\n}\n\nfunction agent<T>(name: string, fn: () => T): T {\n  return tracer.span(\n    `invoke_agent ${name}`,\n    { \"gen_ai.operation.name\": \"invoke_agent\", \"gen_ai.agent.name\": name },\n    fn,\n  );\n}\n\n// MOCK run: a release-notes agent with one subagent.\nconst finalAnswer = agent(\"release-notes\", () => {\n  tracer.span(\n    \"plan release-notes\",\n    { \"gen_ai.operation.name\": \"plan\", \"gen_ai.agent.name\": \"release-notes\" },\n    () => chat(\"mock-large\", 820, 900, 120),\n  );\n  tool(\"read_file\", 30);\n  chat(\"mock-large\", 640, 1400, 210);\n  agent(\"changelog-checker\", () => {\n    chat(\"mock-small\", 210, 600, 40);\n    tool(\"git_log\", 120);\n    chat(\"mock-small\", 180, null, null); // usage not reported yet\n  });\n  try {\n    tool(\"send_email\", 5000, \"timeout\");\n  } catch {\n    // the loop swallows the error and keeps going\n  }\n  chat(\"mock-large\", 450, 1800, 90);\n  return \"Release notes drafted and sent to the team.\";\n});\n```\n\nThree helpers. Three span types.\n\n`chat` records the model and token usage under `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens`.\n\n`tool` records `gen_ai.tool.name`.\n\n`agent` records `gen_ai.agent.name` and wraps everything the agent does.\n\nThe mocked run plans, reads a file, calls a subagent, tries to send an email, and answers.\n\nTwo details are deliberate.\n\nThe second `mock-small` call reports `null` usage. OpenAI's guide says usage can arrive after the turn ends, and that a blank or `null` value \"means the count is unknown. It does not mean the agent used zero tokens.\"\n\nAnd the loop swallows the `send_email` timeout and keeps going. Then the final answer says \"sent.\" That's the bug we're here to catch.\n\n```\nfunction children(id: string | null): Span[] {\n  return tracer.spans.filter((s) => s.parentId === id);\n}\n\nfunction tokens(s: Span): string {\n  const i = s.attrs[\"gen_ai.usage.input_tokens\"];\n  const o = s.attrs[\"gen_ai.usage.output_tokens\"];\n  if (s.attrs[\"gen_ai.operation.name\"] !== \"chat\") return \"\";\n  return i === null || o === null ? \"  tokens=unknown\" : `  in=${i} out=${o}`;\n}\n\nfunction printTree(id: string | null, prefix: string) {\n  const kids = children(id);\n  kids.forEach((s, i) => {\n    const last = i === kids.length - 1;\n    const branch = id === null ? \"\" : prefix + (last ? \"└─ \" : \"├─ \");\n    const ms = `${s.end - s.start}ms`.padStart(7);\n    const label = (branch + s.name).padEnd(40);\n    const err = s.status === \"error\" ? `  error.type=${s.attrs[\"error.type\"]}` : \"\";\n    console.log(`${label} ${s.status.padEnd(5)} ${ms}${tokens(s)}${err}`);\n    printTree(s.spanId, id === null ? \"\" : prefix + (last ? \"   \" : \"│  \"));\n  });\n}\n```\n\nDepth first. Children under parents. Duration, status and tokens on every line.\n\n`tokens()` prints `unknown` instead of `0`.\n\nZero is a number.\n\nUnknown is a different fact.\n\n``` js\nfunction ancestors(s: Span): string[] {\n  const out: string[] = [];\n  let p = tracer.spans.find((x) => x.spanId === s.parentId);\n  while (p) {\n    out.push(p.name);\n    p = tracer.spans.find((x) => x.spanId === p!.parentId);\n  }\n  return out;\n}\n\nfunction failures() {\n  for (const s of tracer.spans.filter((x) => x.status === \"error\")) {\n    console.log(`  ${s.name} (${s.attrs[\"error.type\"]}) <- ${ancestors(s).join(\" <- \")}`);\n  }\n}\n\nfunction claimCheck(answer: string) {\n  const emailFailed = tracer.spans.some(\n    (s) => s.attrs[\"gen_ai.tool.name\"] === \"send_email\" && s.status === \"error\",\n  );\n  if (answer.includes(\"sent\") && emailFailed) {\n    console.log(`  answer says \"sent\", but execute_tool send_email failed`);\n  }\n}\n\nfunction timeByOperation() {\n  const total = new Map<string, number>();\n  for (const s of tracer.spans) {\n    const op = String(s.attrs[\"gen_ai.operation.name\"]);\n    if (op === \"invoke_agent\" || op === \"plan\") continue; // parents, not work\n    total.set(op, (total.get(op) ?? 0) + (s.end - s.start));\n  }\n  for (const [op, ms] of total) console.log(`  ${op.padEnd(13)} ${ms}ms`);\n}\n\nfunction tokensByAgent() {\n  for (const a of tracer.spans.filter((s) => s.attrs[\"gen_ai.operation.name\"] === \"invoke_agent\")) {\n    let known = 0;\n    let unknown = 0;\n    const own = (id: string): Span[] =>\n      children(id).flatMap((c) =>\n        c.attrs[\"gen_ai.operation.name\"] === \"invoke_agent\" ? [] : [c, ...own(c.spanId)],\n      );\n    for (const c of own(a.spanId)) {\n      if (c.attrs[\"gen_ai.operation.name\"] !== \"chat\") continue;\n      const i = c.attrs[\"gen_ai.usage.input_tokens\"];\n      const o = c.attrs[\"gen_ai.usage.output_tokens\"];\n      if (i === null || o === null) unknown++;\n      else known += Number(i) + Number(o);\n    }\n    const note = unknown ? `, ${unknown} chat span unknown` : \"\";\n    console.log(`  ${String(a.attrs[\"gen_ai.agent.name\"]).padEnd(18)} ${known} tokens${note}`);\n  }\n}\n```\n\n`failures()` walks up from every failed span, so you see the error and everything it happened under.\n\n`claimCheck()` is crude on purpose. If the answer says \"sent\" and the email tool failed, it says so. A real version would compare structured claims against tool results. The idea is the same.\n\n`timeByOperation()` sums durations by `gen_ai.operation.name`. It skips `invoke_agent` and `plan`, because those are parents. Counting them would count the same time twice.\n\n`tokensByAgent()` counts each agent's own model calls and stops at subagent boundaries. That matches OpenAI's guide: an agent span's usage covers the agent itself, \"they do not include its subagents.\"\n\n```\nconsole.log(\"Trace (MOCK model, virtual clock)\\n\");\nprintTree(null, \"\");\nconsole.log(`\\nFinal answer: \"${finalAnswer}\"`);\nconsole.log(\"\\nWhat failed, and under what?\");\nfailures();\nconsole.log(\"\\nDoes the answer match the trace?\");\nclaimCheck(finalAnswer);\nconsole.log(\"\\nWhere did the time go?\");\ntimeByOperation();\nconsole.log(\"\\nTokens per agent (subagents counted separately):\");\ntokensByAgent();\n```\n\nRun it:\n\n```\nnpx tsx tracer.ts\n```\n\nYou should see:\n\n```\nTrace (MOCK model, virtual clock)\n\ninvoke_agent release-notes               ok     7450ms\n├─ plan release-notes                    ok      820ms\n│  └─ chat mock-large                    ok      820ms  in=900 out=120\n├─ execute_tool read_file                ok       30ms\n├─ chat mock-large                       ok      640ms  in=1400 out=210\n├─ invoke_agent changelog-checker        ok      510ms\n│  ├─ chat mock-small                    ok      210ms  in=600 out=40\n│  ├─ execute_tool git_log               ok      120ms\n│  └─ chat mock-small                    ok      180ms  tokens=unknown\n├─ execute_tool send_email               error  5000ms  error.type=timeout\n└─ chat mock-large                       ok      450ms  in=1800 out=90\n\nFinal answer: \"Release notes drafted and sent to the team.\"\n\nWhat failed, and under what?\n  execute_tool send_email (timeout) <- invoke_agent release-notes\n\nDoes the answer match the trace?\n  answer says \"sent\", but execute_tool send_email failed\n\nWhere did the time go?\n  chat          2300ms\n  execute_tool  5150ms\n\nTokens per agent (subagents counted separately):\n  release-notes      4520 tokens\n  changelog-checker  640 tokens, 1 chat span unknown\n```\n\nThe final answer is confident.\n\nThe trace disagrees.\n\n`send_email` ran for 5000ms and ended in `error.type=timeout`. The root agent span still says `ok`, because the loop caught the error. A run-level status would have told you nothing.\n\nThat's the lesson from OpenAI's monitor, in miniature. A failed attempt is still an attempt. The span records that it happened, not just whether it worked.\n\nThe time question has a boring answer, which is the best kind. 5150 of 7450ms went to tool calls, and 5000 of those were one email that never went out.\n\nAnd the subagent's tokens come with a footnote. 640 known, 1 call unknown. Not 640. Not zero. At least 640.\n\n**The answer is what the agent says. The trace is what the agent did.**\n\nThis is a teaching tracer. Here is what a real one would need.\n\nWe keep spans in an array. Real systems use an OpenTelemetry SDK with context propagation, sampling and an OTLP exporter. The span names and attributes are the part worth copying.\n\nA stack works for synchronous code. Two subagents running at once need real context propagation, or every span ends up under whichever one started last.\n\nWe don't record prompts or tool arguments. The conventions treat message content as opt-in for a reason. Traces of agents that read inboxes are full of other people's data.\n\nOpenAI's report says DNS activity was logged, but an infrastructure detector for anomalous DNS excluded the affected environment. A trace nobody queries is a very detailed diary. Our four questions are hard-coded. Real ones belong in alerts that can stop a run.\n\nEverything `gen_ai.*` here is Development status. Two changes landed this week alone. Pin a version and expect renames.\n\nLogs answer: what happened?\n\nTraces answer: what happened, under what, for how long, and at whose cost?\n\nFor a chatbot, the transcript was enough.\n\nFor an agent with tools, subagents and a budget, it isn't.\n\n```\n┌──────────────────────────────────────────┐\n│ invoke_agent                             │\n│   ├── plan ──→ chat                      │\n│   ├── execute_tool                       │\n│   ├── invoke_agent (subagent)            │\n│   │     ├── chat                         │\n│   │     └── execute_tool                 │\n│   └── chat ──→ final answer (a claim)    │\n└──────────────────────────────────────────┘\n                     ↓\n        failures · claims · time · tokens\n```\n\nThe span provides a unit of work.\n\nThe parent link provides structure.\n\nThe status provides the truth about each step.\n\nThe attributes provide a shared vocabulary.\n\nThe clock provides cost in time.\n\nThe usage provides cost in tokens, or an honest \"unknown.\"\n\nThe final answer provides a claim to check.\n\n**If your agent can't show its work, you're grading its confidence.**\n\nI'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access. They work inside a lane, leave a record of what they did, and ask before doing anything you'd want to see first.\n\nIf the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.", "url": "https://wpnews.pro/news/stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript", "canonical_source": "https://dev.to/bobbyhalljr/stop-trusting-your-agents-final-answer-build-a-tiny-agent-tracer-in-typescript-24dl", "published_at": "2026-10-04 17:07:15+00:00", "updated_at": "2026-10-04 17:12:43.900389+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "mlops", "ai-tools"], "entities": ["OpenAI", "Sam Altman", "TypeScript", "OpenTelemetry", "tiny-agent-tracer", "Roster", "Node.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript", "markdown": "https://wpnews.pro/news/stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript.md", "text": "https://wpnews.pro/news/stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript.txt", "jsonld": "https://wpnews.pro/news/stop-trusting-your-agent-s-final-answer-build-a-tiny-agent-tracer-in-typescript.jsonld"}}