On September 20, an OpenAI research agent was asked to identify the author of a blog post.
It ended the run politely.
"I couldn't reliably establish" who it was, it told the user (in OpenAI's English translation). It asked for the original wording, or the title.
Humble. Helpful. Nothing to see.
The trace told a different story.
OpenAI has d training, evaluation and inference with tool use for its most capable models while it closes the gap. In the post announcing the reports site, Sam Altman said the company is working to understand "petabytes of agent activity logs."
So here is my contrarian take.
The final answer is the least trustworthy thing an agent produces.
Not because agents lie.
Because the answer is a summary, written by the same system you're trying to check.
The answer is a claim. The trace is the evidence.
The tooling world clearly agrees. Look at the last two weeks:
gen_ai.skill.* attributes for the execute tool span (#498) and a gen_ai.main_agent entity (#270). The span names are already there: invoke_agent {gen_ai.agent.name}, execute_tool {gen_ai.tool.name}, {gen_ai.operation.name} {gen_ai.request.model} for a model call, and a plan span since May. All of it is still marked Development.
Everyone is shipping tools to read traces.
Which only helps if your agent writes a good one.
In Evals for Agents: Did It Stay in Scope? I graded runs after the fact. This post is about recording the run well enough that there's something to grade.
Let's build a tiny tracer.
By the end, you'll run one command:
npx tsx tracer.ts
And watch a mocked agent claim it sent an email that, according to its own trace, timed out.
Smaller stakes than a DNS side channel. Same shape: the answer and the trace disagree.
No API key.
No OpenTelemetry SDK.
Just TypeScript, and the same span names the conventions use.
One honesty note: this is not how OpenAI, AWS or anyone else records traces. It's my small model of the shape their docs describe.
Code: github.com/bobbyhalljr/tiny-agent-tracer
One run.
One root span for the agent.
Children for planning, model calls, tool calls and a subagent.
Every span gets a name, a parent, a start, an end, a status and a few gen_ai.* attributes.
Then we ask the trace four questions:
It's also a small version of something I care about in Roster: an AI employee should leave evidence, not just a confident summary.
The model, the tools and every duration are a mock on a virtual clock, so the output is identical on every run.
You will need Node.js 18 or newer.
mkdir tiny-agent-tracer
cd tiny-agent-tracer
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following blocks, in order, as tracer.ts.
// tiny-agent-tracer: trace one agent run with OpenTelemetry GenAI span names.
// The model, the tools and every duration are a MOCK on a virtual clock, so
// the output is the same on every run. No network. No API key.
type Attr = string | number | boolean | null;
type Span = {
spanId: string;
parentId: string | null;
name: string;
attrs: Record<string, Attr>;
start: number;
end: number;
status: "ok" | "error";
};
A span is a unit of work with a parent.
That's the whole trick.
The parent link turns a flat log into a tree. The tree is what tells you that a tool call belonged to a subagent and not to the root.
Attr allows null on purpose. We'll need it for tokens.
class Tracer {
spans: Span[] = [];
private stack: Span[] = [];
private now = 0;
private nextId = 1;
advance(ms: number) {
this.now += ms;
}
span<T>(name: string, attrs: Record<string, Attr>, fn: (s: Span) => T): T {
const parent = this.stack[this.stack.length - 1];
const s: Span = {
spanId: `s${this.nextId++}`,
parentId: parent ? parent.spanId : null,
name,
attrs,
start: this.now,
end: this.now,
status: "ok",
};
this.spans.push(s);
this.stack.push(s);
try {
return fn(s);
} catch (err) {
s.status = "error";
s.attrs["error.type"] = (err as Error).message;
throw err;
} finally {
s.end = this.now;
this.stack.pop();
}
}
}
span() pushes a span on a stack, runs your function, and pops it.
Whatever runs inside becomes a child.
If the function throws, the span is marked error, gets an error.type, and the error keeps going. The tracer records failures. It doesn't hide them.
The clock is fake. advance(ms) moves time forward. Real tracers read a real clock, but a fake one makes the output deterministic, which is nice for a tutorial and essential for a test.
A trace you can't reproduce is a screenshot.
const tracer = new Tracer();
function chat(model: string, ms: number, input: number | null, output: number | null) {
return tracer.span(
`chat ${model}`,
{
"gen_ai.operation.name": "chat",
"gen_ai.request.model": model,
"gen_ai.usage.input_tokens": input,
"gen_ai.usage.output_tokens": output,
},
() => tracer.advance(ms),
);
}
function tool(name: string, ms: number, fail?: string) {
return tracer.span(
`execute_tool ${name}`,
{ "gen_ai.operation.name": "execute_tool", "gen_ai.tool.name": name },
() => {
tracer.advance(ms);
if (fail) throw new Error(fail);
},
);
}
function agent<T>(name: string, fn: () => T): T {
return tracer.span(
`invoke_agent ${name}`,
{ "gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": name },
fn,
);
}
// MOCK run: a release-notes agent with one subagent.
const finalAnswer = agent("release-notes", () => {
tracer.span(
"plan release-notes",
{ "gen_ai.operation.name": "plan", "gen_ai.agent.name": "release-notes" },
() => chat("mock-large", 820, 900, 120),
);
tool("read_file", 30);
chat("mock-large", 640, 1400, 210);
agent("changelog-checker", () => {
chat("mock-small", 210, 600, 40);
tool("git_log", 120);
chat("mock-small", 180, null, null); // usage not reported yet
});
try {
tool("send_email", 5000, "timeout");
} catch {
// the loop swallows the error and keeps going
}
chat("mock-large", 450, 1800, 90);
return "Release notes drafted and sent to the team.";
});
Three helpers. Three span types.
chat records the model and token usage under gen_ai.usage.input_tokens and gen_ai.usage.output_tokens.
tool records gen_ai.tool.name.
agent records gen_ai.agent.name and wraps everything the agent does.
The mocked run plans, reads a file, calls a subagent, tries to send an email, and answers.
Two details are deliberate.
The second mock-small call reports null usage. OpenAI's guide says usage can arrive after the turn ends, and that a blank or null value "means the count is unknown. It does not mean the agent used zero tokens."
And the loop swallows the send_email timeout and keeps going. Then the final answer says "sent." That's the bug we're here to catch.
function children(id: string | null): Span[] {
return tracer.spans.filter((s) => s.parentId === id);
}
function tokens(s: Span): string {
const i = s.attrs["gen_ai.usage.input_tokens"];
const o = s.attrs["gen_ai.usage.output_tokens"];
if (s.attrs["gen_ai.operation.name"] !== "chat") return "";
return i === null || o === null ? " tokens=unknown" : ` in=${i} out=${o}`;
}
function printTree(id: string | null, prefix: string) {
const kids = children(id);
kids.forEach((s, i) => {
const last = i === kids.length - 1;
const branch = id === null ? "" : prefix + (last ? "ββ " : "ββ ");
const ms = `${s.end - s.start}ms`.padStart(7);
const label = (branch + s.name).padEnd(40);
const err = s.status === "error" ? ` error.type=${s.attrs["error.type"]}` : "";
console.log(`${label} ${s.status.padEnd(5)} ${ms}${tokens(s)}${err}`);
printTree(s.spanId, id === null ? "" : prefix + (last ? " " : "β "));
});
}
Depth first. Children under parents. Duration, status and tokens on every line.
tokens() prints unknown instead of 0.
Zero is a number.
Unknown is a different fact.
function ancestors(s: Span): string[] {
const out: string[] = [];
let p = tracer.spans.find((x) => x.spanId === s.parentId);
while (p) {
out.push(p.name);
p = tracer.spans.find((x) => x.spanId === p!.parentId);
}
return out;
}
function failures() {
for (const s of tracer.spans.filter((x) => x.status === "error")) {
console.log(` ${s.name} (${s.attrs["error.type"]}) <- ${ancestors(s).join(" <- ")}`);
}
}
function claimCheck(answer: string) {
const emailFailed = tracer.spans.some(
(s) => s.attrs["gen_ai.tool.name"] === "send_email" && s.status === "error",
);
if (answer.includes("sent") && emailFailed) {
console.log(` answer says "sent", but execute_tool send_email failed`);
}
}
function timeByOperation() {
const total = new Map<string, number>();
for (const s of tracer.spans) {
const op = String(s.attrs["gen_ai.operation.name"]);
if (op === "invoke_agent" || op === "plan") continue; // parents, not work
total.set(op, (total.get(op) ?? 0) + (s.end - s.start));
}
for (const [op, ms] of total) console.log(` ${op.padEnd(13)} ${ms}ms`);
}
function tokensByAgent() {
for (const a of tracer.spans.filter((s) => s.attrs["gen_ai.operation.name"] === "invoke_agent")) {
let known = 0;
let unknown = 0;
const own = (id: string): Span[] =>
children(id).flatMap((c) =>
c.attrs["gen_ai.operation.name"] === "invoke_agent" ? [] : [c, ...own(c.spanId)],
);
for (const c of own(a.spanId)) {
if (c.attrs["gen_ai.operation.name"] !== "chat") continue;
const i = c.attrs["gen_ai.usage.input_tokens"];
const o = c.attrs["gen_ai.usage.output_tokens"];
if (i === null || o === null) unknown++;
else known += Number(i) + Number(o);
}
const note = unknown ? `, ${unknown} chat span unknown` : "";
console.log(` ${String(a.attrs["gen_ai.agent.name"]).padEnd(18)} ${known} tokens${note}`);
}
}
failures() walks up from every failed span, so you see the error and everything it happened under.
claimCheck() is crude on purpose. If the answer says "sent" and the email tool failed, it says so. A real version would compare structured claims against tool results. The idea is the same.
timeByOperation() sums durations by gen_ai.operation.name. It skips invoke_agent and plan, because those are parents. Counting them would count the same time twice.
tokensByAgent() counts each agent's own model calls and stops at subagent boundaries. That matches OpenAI's guide: an agent span's usage covers the agent itself, "they do not include its subagents."
console.log("Trace (MOCK model, virtual clock)\n");
printTree(null, "");
console.log(`\nFinal answer: "${finalAnswer}"`);
console.log("\nWhat failed, and under what?");
failures();
console.log("\nDoes the answer match the trace?");
claimCheck(finalAnswer);
console.log("\nWhere did the time go?");
timeByOperation();
console.log("\nTokens per agent (subagents counted separately):");
tokensByAgent();
Run it:
npx tsx tracer.ts
You should see:
Trace (MOCK model, virtual clock)
invoke_agent release-notes ok 7450ms
ββ plan release-notes ok 820ms
β ββ chat mock-large ok 820ms in=900 out=120
ββ execute_tool read_file ok 30ms
ββ chat mock-large ok 640ms in=1400 out=210
ββ invoke_agent changelog-checker ok 510ms
β ββ chat mock-small ok 210ms in=600 out=40
β ββ execute_tool git_log ok 120ms
β ββ chat mock-small ok 180ms tokens=unknown
ββ execute_tool send_email error 5000ms error.type=timeout
ββ chat mock-large ok 450ms in=1800 out=90
Final answer: "Release notes drafted and sent to the team."
What failed, and under what?
execute_tool send_email (timeout) <- invoke_agent release-notes
Does the answer match the trace?
answer says "sent", but execute_tool send_email failed
Where did the time go?
chat 2300ms
execute_tool 5150ms
Tokens per agent (subagents counted separately):
release-notes 4520 tokens
changelog-checker 640 tokens, 1 chat span unknown
The final answer is confident.
The trace disagrees.
send_email ran for 5000ms and ended in error.type=timeout. The root agent span still says ok, because the loop caught the error. A run-level status would have told you nothing.
That's the lesson from OpenAI's monitor, in miniature. A failed attempt is still an attempt. The span records that it happened, not just whether it worked.
The time question has a boring answer, which is the best kind. 5150 of 7450ms went to tool calls, and 5000 of those were one email that never went out.
And the subagent's tokens come with a footnote. 640 known, 1 call unknown. Not 640. Not zero. At least 640.
The answer is what the agent says. The trace is what the agent did.
This is a teaching tracer. Here is what a real one would need.
We keep spans in an array. Real systems use an OpenTelemetry SDK with context propagation, sampling and an OTLP exporter. The span names and attributes are the part worth copying.
A stack works for synchronous code. Two subagents running at once need real context propagation, or every span ends up under whichever one started last.
We don't record prompts or tool arguments. The conventions treat message content as opt-in for a reason. Traces of agents that read inboxes are full of other people's data.
OpenAI's report says DNS activity was logged, but an infrastructure detector for anomalous DNS excluded the affected environment. A trace nobody queries is a very detailed diary. Our four questions are hard-coded. Real ones belong in alerts that can stop a run.
Everything gen_ai.* here is Development status. Two changes landed this week alone. Pin a version and expect renames.
Logs answer: what happened?
Traces answer: what happened, under what, for how long, and at whose cost?
For a chatbot, the transcript was enough.
For an agent with tools, subagents and a budget, it isn't.
ββββββββββββββββββββββββββββββββββββββββββββ
β invoke_agent β
β βββ plan βββ chat β
β βββ execute_tool β
β βββ invoke_agent (subagent) β
β β βββ chat β
β β βββ execute_tool β
β βββ chat βββ final answer (a claim) β
ββββββββββββββββββββββββββββββββββββββββββββ
β
failures Β· claims Β· time Β· tokens
The span provides a unit of work.
The parent link provides structure.
The status provides the truth about each step.
The attributes provide a shared vocabulary.
The clock provides cost in time.
The usage provides cost in tokens, or an honest "unknown."
The final answer provides a claim to check.
If your agent can't show its work, you're grading its confidence.
I'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access. They work inside a lane, leave a record of what they did, and ask before doing anything you'd want to see first.
If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.