cd /news/developer-tools/a-cli-that-shows-which-sub-agent-bur… Β· home β€Ί topics β€Ί developer-tools β€Ί article
[ARTICLE Β· art-78098] src=github.com β†— pub= topic=developer-tools verified=true sentiment=Β· neutral

A CLI that shows which sub-agent burned your tokens in one run

A new command-line tool called token-trace-viewer lets developers trace token and dollar costs across sub-agents in a single multi-agent run by reading OpenInference-style JSONL trace files. The Python 3.10+ tool, which requires no server, dashboard, account, or pip install, outputs plain tables showing which sub-agent burned the most tokens, which step cost the most money, and where the tool failed. It includes sample traces and supports flags for JSON output, agent/model filtering, top-N rows, and custom pricing files.

read11 min views1 publishedJul 29, 2026
A CLI that shows which sub-agent burned your tokens in one run
Image: source

Explain where the tokens (and dollars) went in a single agent run.

token-trace-viewer

reads an OpenInference-style JSONL trace produced by one multi-agent run and answers, from the command line, in plain tables:

  • Which sub-agent burned the most tokens, and on which step?
  • Which single step cost the most money?
  • How much of the file did the tool actually understand, and where did it fail?

No server, no dashboard, no account. Point it at a log file, get a table back.

$ python3 cost.py sample.jsonl
 1  planner     verify_output        15,240    430  $0.0869       63.0%
 2  planner     decompose_goal        2,610    388  $0.0227       16.5%
 3  planner     plan_task             1,840    512  $0.0220       15.9%
 4  planner     route_to_researcher     940     64  $0.0063        4.6%
 5  researcher  search_web            4,200    310      n/a         n/a
 6  researcher  web_search_api            0      0      n/a         n/a
 ...

note: 8 step(s) have no price for their model (or no llm.model_name at all)
and are shown with cost "n/a"; they are excluded from % of run total.

covered 12 of 12 lines; 0 unreadable, 0 with no agent name, 0 duplicate span(s) dropped
  • Python 3.10+(usesX | None

union type syntax in dataclass fields) - Standard library only β€” no pip install

needed

python3 report.py <path-to-log.jsonl>    # token totals by agent, then by step
python3 cost.py <path-to-log.jsonl>      # every step ranked by dollar cost
python3 repeats.py <path-to-log.jsonl>   # context blocks resent across steps
python3 parser.py <path-to-log.jsonl>    # quick parse summary (lines parsed/skipped)

python3 tokentrace.py report <path-to-log.jsonl>
python3 tokentrace.py cost <path-to-log.jsonl>
python3 tokentrace.py repeats <path-to-log.jsonl>

Useful flags on report.py

/ cost.py

/ repeats.py

(and their tokentrace.py

equivalents):

Flag Effect
--json
Print machine-readable JSON instead of a table
--agent NAME
Only include steps from this agent (report.py , cost.py )
--model NAME
Only include steps from this model (report.py , cost.py )
--top N
Only show the top N rows
--prices FILE
Use a pricing file other than prices.json (cost.py , repeats.py )

Two sample traces are included:

sample.jsonl

β€” a clean 12-step, 3-agent (planner β†’ researcher β†’ writer) research pipelinesample-broken.jsonl

β€” the same trace with real-world damage injected (a truncated line, a stray unescaped quote, a bare blank line, a span with noagent.name

) to demonstrate the tool's error handling

python3 report.py sample.jsonl
python3 cost.py sample-broken.jsonl
parser.py             shared JSONL β†’ Step parsing (stdlib only, no cost/ranking logic)
report.py             token totals by agent β†’ step
cost.py               cost ranking by step
repeats.py            context blocks resent more than once across steps
tokentrace.py          single CLI entry point over report/cost/repeats/parse
prices.json.example    sample pricing-overrides file (copy to prices.json to use)
tests/                 stdlib unittest suite for all of the above
sample.jsonl           clean 3-agent example trace
sample-broken.jsonl    the same trace with real-world corruption injected
scope.md               what this tool does and deliberately does not do, and why
existence-check.md     survey of existing tools that informed the scope
python3 -m unittest discover -s tests -v

Stdlib unittest

, no extra dependencies. Covers parser.py

(malformed lines, negative/null token counts, duplicate span_id

handling, timestamp parsing), cost.py

(pricing lookup, sorting, prices.json

overrides), report.py

(aggregation and rendering), and repeats.py

(repeat detection and the wasted-token estimate) β€” including full runs against both sample.jsonl

and sample-broken.jsonl

.

Trace dashboards like LangSmith, Langfuse, Helicone, and Arize Phoenix all answer cross-run questions well β€” cost trends over a week, spend per project, per model. None of them make it easy to answer a much narrower question an engineer actually asks after a spike: "why did this one run cost $4.20, and which sub-agent was responsible?"

A short survey of five tools (existence-check.md

) found:

Tool Per-sub-agent cost ranking for one run? Note
LangSmith Dashboard-level only Custom widget, aggregates across runs, no single-run breakdown
Langfuse Dashboard-level only Same shape as LangSmith
Helicone No (tags-only) Requires you to tag every call and write your own SQL
Phoenix / Arize No Cost shown per trace/span/project; no grouping or ranking
OpenInference spec N/A Defines the fields (agent.name , llm.token_count.* , ...) but ships no UI

token-trace-viewer

reads exactly those OpenInference fields and turns one trace file into a per-run, per-agent, per-step breakdown β€” nothing more. See scope.md

for the full rationale and out-of-scope list.

Script Answers Sorted by
report.py
How many input/output tokens did each agent β€” and each step within that agent β€” use? Token volume, biggest agent first, biggest step within each agent
cost.py
What did each individual step cost in dollars, across the whole run? Dollar cost, highest first
repeats.py
Which context blocks (system/user/assistant messages) were sent more than once across steps, and roughly how much did that waste? Estimated wasted tokens, highest first
parser.py
(library + CLI smoke test) Parse the log and report how many lines were understood β€”
tokentrace.py
One entry point for the three commands above (`tokentrace.py report cost repeats <file>` )
β€”

report.py

, cost.py

, and repeats.py

are all built on the same shared parser (parser.py

) and all print a coverage line at the end so you always know how much of the file the tool actually read. Each also accepts --json

for machine-readable output, and --agent

/--model

/--top

(where applicable) to filter or limit a large trace.

No dashboard, no server. One command, text output, nothing hosted.No trends across runs. It explains a single trace file, not a fleet of runs over time β€” that's what LangSmith/Langfuse are for.No optimization advice. It tells you a step cost $0.09 and ranks it; it doesn't tell you how to make it cheaper.No instrumentation. It doesn't wrap your LLM calls or generate traces β€” you bring a log, it reads it.One provider's prices, hardcoded. SeePricingbelow.

writer                   37,810 in  5,365 out
    emit_answer          16,010 in  1,120 out
    revise_draft         12,680 in  1,905 out
    draft_section         9,120 in  2,340 out
researcher               34,380 in  2,428 out
    summarize_findings   18,430 in  1,276 out
    read_results         11,750 in    842 out
    search_web            4,200 in    310 out
    web_search_api            0 in      0 out
    fetch_page                0 in      0 out
planner                  20,630 in  1,394 out
    verify_output        15,240 in    430 out
    decompose_goal        2,610 in    388 out
    plan_task             1,840 in    512 out
    route_to_researcher     940 in     64 out
TOTAL                    92,820 in  9,187 out
covered 12 of 12 lines; 0 unreadable, 0 with no agent name, 0 duplicate span(s) dropped

Agents are sorted by total tokens (input + output), highest first. Steps within each agent are sorted the same way. A step name that repeats within an agent (e.g. a retry loop) is folded into one row and annotated with (x<count>)

.

Priced rows sort to the top by cost, highest first; unpriced rows (no known price for the model, or no llm.model_name

at all) sink to the bottom, marked n/a

, and are excluded from the % OF TOTAL

column so the percentages still add up to 100% of priced spend.

The gap noted in scope.md

: no surveyed tool (see existence-check.md

) flags identical context resent across steps in a single run. repeats.py

hashes every system/user/assistant message body in the trace and flags any hash that shows up under more than one step:

$ python3 repeats.py sample.jsonl
HASH      ROLE    SENDS  CHARS  EST WASTED TOKENS  EST WASTED $  SEEN IN
d8c5cd4f  system      9  3,066              6,128           n/a  planner/plan_task, planner/decompose_goal, researcher/search_web, +6 more

note: token/cost figures are estimated at ~4 chars per token (traces carry per-step totals, not
per-message token counts), and only cover blocks priced in the model's rate table; 'n/a' means at
least one resend used an unpriced model.

covered 12 of 12 lines; 0 unreadable, 0 with no agent name, 0 duplicate span(s) dropped

Here, the same ~3KB system prompt was sent unchanged on 9 of the run's 9 LLM calls β€” 8 of those sends were pure repeats of the first, at an estimated ~6,100 wasted input tokens.

Two things this deliberately does not do, to stay honest about what a trace actually contains:

  • It hashes message content rather than storing it, so a trace with megabytes of prompt text doesn't need to be held in memory just to compare it. You get a hash prefix and a length, not a diff.
  • OpenInference spans carry a per-steptoken total, not aper-messageone, so the token cost of one resent block can't be read off the trace directly β€” it's estimated from character length at roughly 4 characters per token, and the output says so every time.
$ python3 parser.py sample.jsonl
parsed 12 steps, skipped 0 bad lines, 0 step(s) had no agent name, 0 duplicate span(s) dropped

report.py

, cost.py

, and repeats.py

all import parse_log()

from this module rather than re-implementing parsing.

Real traces come from streaming exporters: processes get killed mid-write, strings arrive unescaped, lines show up empty. parser.py

reads the file one line at a time β€” never json.load()

s the whole thing β€” so:

  • a truncated or malformed line is skipped and counted, not fatal to the run - a syntactically valid JSON value that isn't an object (a bare string, number, or list) is treated the same way
  • a line that parses fine but is missing agent.name

is counted separately as**"unknown agent"**β€” aquietfailure mode where the run looks healthy but the tool actually learned nothing from that line, so it's surfaced instead of hidden - token counts that arrive as null

, missing, a numeric string, ornegative(exporter corruption, not a real count) are coerced to0

instead of crashing the run or skewing totals - a span whose span_id

has already been seen is aduplicate, not a new step β€” a common artifact of a streaming exporter that crashed and re-emitted a span on restart. It is dropped and counted separately rather than silently double-counted into the totals

Every run ends with a coverage line β€” covered X of Y lines; N unreadable, M with no agent name, K duplicate span(s) dropped

β€” so you always know how much of the file the numbers above actually represent.

Running report.py

against the intentionally-damaged sample-broken.jsonl

shows this in action:

$ python3 report.py sample-broken.jsonl
...
unknown                       0 in      0 out
    route_to_researcher       0 in      0 out
TOTAL                    73,260 in  7,615 out
covered 8 of 12 lines; 4 unreadable, 1 with no agent name, 0 duplicate span(s) dropped

If a span carries start_time

/end_time

timestamps, report.py

and cost.py

also print a run duration: N.Ns

line β€” the span from the earliest start to the latest end seen in the trace. It's omitted when the trace has fewer than two distinct timestamps to measure a duration from (as in the two sample files above, which only carry one incidental timestamp between them).

cost.py

ships one hardcoded default price table (DEFAULT_PRICES

in cost.py

), sourced from the Anthropic pricing page and checked on the date noted in the file:

Model Input $ / 1M tokens Output $ / 1M tokens
claude-opus-5-0
$5.00 $25.00

Any step whose llm.model_name

isn't in that table β€” or has no model name at all (e.g. some CHAIN

/TOOL

spans) β€” is priced as n/a

rather than guessed at, and excluded from the percentage-of-total column.

To price additional models or providers without editing code, drop a prices.json

next to cost.py

:

{
  "claude-opus-5-0": { "input": 5.0, "output": 25.0 },
  "some-other-model": { "input": 1.5, "output": 6.0 }
}

Entries there are overlaid on top of the built-in defaults (same model name overrides the price; new model names are added). repeats.py

uses the same table, since it estimates the dollar cost of a resent block, too. Use --prices <file>

to point at a pricing file somewhere other than the default location.

Each line of the input file is one JSON object (a span), following the OpenInference semantic conventions. The fields this tool reads:

Field Used for
attributes["agent.name"]
Grouping (defaults to "unknown" if absent)
attributes["graph.node.name"] (falls back to top-level "name" )
Step label
attributes["llm.model_name"]
Cost lookup
attributes["llm.token_count.prompt"]
Input tokens
attributes["llm.token_count.completion"]
Output tokens
attributes["llm.input_messages.<i>.message.role"] / ...content"]
Repeated-context detection (repeats.py ); content is hashed, not stored
top-level "span_id"
De-duplicating spans re-emitted by a restarted exporter
top-level "start_time" / "end_time"
The run duration line in report.py / cost.py , when present

Tool names and any other span fields not listed above are still ignored.

Per scope.md

:

Repeated-context detectionβ€” shipped asrepeats.py

.- Prices for additional model providers beyond the single hardcoded entry β€” the prices.json

override file (seePricing) makes this possible without code changes, but no additional providers are pre-populated. - Framework-specific log formats (only OpenInference JSONL is read today).

No license file is currently included in this repository.

── more in #developer-tools 4 stories Β· sorted by recency
── more on @token-trace-viewer 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/a-cli-that-shows-whi…] indexed:0 read:11min 2026-07-29 Β· β€”