Stop Trusting Your Agent's Final Answer: Build a Tiny Agent Tracer in TypeScript A developer published a TypeScript tutorial and open-source project, tiny-agent-tracer, that records an agent run as a tree of OpenTelemetry-style spans so the execution trace can be compared against the agent's final answer. The tracer uses mocked models, tools and a virtual clock to produce deterministic output with no API key or OpenTelemetry SDK, and demonstrates a mocked agent claiming it sent an email that its own trace shows timed out. On September 20, an OpenAI research agent was asked to identify the author of a blog post. It ended the run politely. "I couldn't reliably establish" who it was, it told the user in OpenAI's English translation . It asked for the original wording, or the title. Humble. Helpful. Nothing to see. The trace told a different story. OpenAI has paused training, evaluation and inference with tool use for its most capable models while it closes the gap. In the post announcing the reports site https://techcrunch.com/2026/09/28/openai-still-doesnt-seem-to-have-a-handle-on-all-of-its-rogue-ai-activity/ , Sam Altman said the company is working to understand "petabytes of agent activity logs." So here is my contrarian take. The final answer is the least trustworthy thing an agent produces. Not because agents lie. Because the answer is a summary, written by the same system you're trying to check. The answer is a claim. The trace is the evidence. The tooling world clearly agrees. Look at the last two weeks: gen ai.skill. attributes for the execute tool span 498 and a gen ai.main agent entity 270 . The span names are already there: invoke agent {gen ai.agent.name} , execute tool {gen ai.tool.name} , {gen ai.operation.name} {gen ai.request.model} for a model call, and a plan span since May. All of it is still marked Development. Everyone is shipping tools to read traces. Which only helps if your agent writes a good one. In Evals for Agents: Did It Stay in Scope? https://dev.to/bobbyhalljr/evals-for-agents-did-it-stay-in-scope-build-a-tiny-one-in-typescript-50j2 I graded runs after the fact. This post is about recording the run well enough that there's something to grade. Let's build a tiny tracer. By the end, you'll run one command: npx tsx tracer.ts And watch a mocked agent claim it sent an email that, according to its own trace, timed out. Smaller stakes than a DNS side channel. Same shape: the answer and the trace disagree. No API key. No OpenTelemetry SDK. Just TypeScript, and the same span names the conventions use. One honesty note: this is not how OpenAI, AWS or anyone else records traces. It's my small model of the shape their docs describe. Code: github.com/bobbyhalljr/tiny-agent-tracer https://github.com/bobbyhalljr/tiny-agent-tracer One run. One root span for the agent. Children for planning, model calls, tool calls and a subagent. Every span gets a name, a parent, a start, an end, a status and a few gen ai. attributes. Then we ask the trace four questions: It's also a small version of something I care about in Roster https://get-roster.com : an AI employee should leave evidence, not just a confident summary. The model, the tools and every duration are a mock on a virtual clock, so the output is identical on every run. You will need Node.js 18 or newer. mkdir tiny-agent-tracer cd tiny-agent-tracer npm init -y npm install --save-dev typescript tsx @types/node Save the following blocks, in order, as tracer.ts . // tiny-agent-tracer: trace one agent run with OpenTelemetry GenAI span names. // The model, the tools and every duration are a MOCK on a virtual clock, so // the output is the same on every run. No network. No API key. type Attr = string | number | boolean | null; type Span = { spanId: string; parentId: string | null; name: string; attrs: Record