# Your Team’s AI Spend Is a Black Box — Even to You

> Source: <https://dev.to/alpomar/your-teams-ai-spend-is-a-black-box-even-to-you-l60>
> Published: 2026-09-29 13:03:00+00:00

“I ran out of tokens.”

“Me too. Can we get more?”

“Sure,” says the business. “Where did they go?”

Silence. Shrugs. Someone mumbles something about refactoring.

That exchange is happening in engineering orgs everywhere right now. You’re under pressure to control the bill, not just explain it — and you can’t control what you can’t see.

We know the total bill. We don’t know which feature it paid for.

That’s not a tooling gap. It’s a visibility gap. And it’s costing you two things at once: a straight answer on return for the spend, and any way to weigh cost against value at all.

Cost attribution means tagging assistant usage to the piece of work it went into. Not a team. Not a month. A feature, a ticket, one line in a backlog.

It’s the half of the equation nobody has. A rough number exists for engineering time (e.g., person-days x daily $), but it’s never included what the assistant itself cost. What most teams have instead is a dashboard showing total tokens per person, maybe per team, refreshed monthly — useful for a finance conversation about the subscription, useless for putting a cost next to any one feature’s value.

That’s not an oversight. No assistant vendor has much incentive to make their own token spend easy to scrutinise.

Once you have cost attribution, you can finally put the two halves next to each other: this feature cost this much, and here’s what we believed it was worth. Sometimes they’ll agree. Sometimes the gap will be the most useful thing in the room.

Here’s where the same data becomes your leverage in the conversation with product.

Two tickets, same sprint. PROJ-118 burned 400,000 tokens and shipped a feature two enterprise customers had asked for by name. PROJ-204 burned 2.6 million (six times the spend) fixing a papercut nobody had complained about.

Without attribution, both look identical in the backlog: five story points each, both “done.” With attribution, they’re not the same at all — one delivered named value cheaply, the other cost six times as much for a fix nobody asked for.

That’s the evidence you bring into the prioritisation conversation yourself: proof this sprint’s “quick win” has a track record of being anything but.

The good news: all of this is already sitting on your engineers’ machines, waiting to be read.

A Claude Code session isn’t just recorded somewhere — the session is the transcript. It writes itself to disk as it goes, a plain JSONL file sitting locally on the engineer’s machine: every prompt, every tool call, every token count. Nobody has to switch it on. It’s already happening.

The assistant also fires events at named moments in a session’s life — SessionStart when a session opens, SessionEnd when it closes, and others besides, though those two are all this needs. A hook is what you register to catch one: nothing more exotic than a line in a config file pointing an event at a script of your choosing.

```
on SessionStart  → run reminder_script
on SessionEnd    → run export_script
```

That’s the whole mechanism. No agent to install, no proxy to sit in front of the assistant, no changes to how engineers work. You’re reading a file that already exists, at two moments that already happen.

With that in mind, the shape of this is simpler than it sounds. What follows is pseudo-code, meant to build intuition for how the three pieces fit together. If you’d rather start from real code, see the reference repo on GitHub.

**1. A feature ID reminder, at SessionStart.** Ask for the feature ID the session relates to, and check again a few prompts later if it never showed up:

on session_start:

    ask("which ticket is this session for?")

on each prompt:

    if prompt_count in [2, 5, 10, 20] and no ticket_id seen yet:

        ask_again()

**2. An export, at SessionEnd.** When the engineer exits Claude Code, it runs whatever script settings.json points it to, passing that script the session's details. Your script's only job is to copy that transcript somewhere durable, off the engineer's machine, without making them wait for it:

```
on session_end(session_info):
    transcript = session_info.transcript_path
    spawn_background(copy transcript → database)
    return immediately
```

**3. A batch job, run on a schedule.** Step 2 only ever pushed raw transcripts into durable storage — nothing has parsed them yet. This job walks every transcript sitting there unprocessed, and turns each one into a single row of metrics in a separate table:

```
for each raw session not yet in the metrics table:
    tokens = sum_by_type(session.events)
    feature_id = find_ticket_reference(session.raw_text)
    cost = weighted(tokens, session.model)
    duration = session.end_time - session.start_time
    write_row(feature_id, tokens, cost, model, duration)
```

Token types aren’t priced the same, and neither are models, so weighted() needs both:

```
weighted(tokens, model):
    ratio = (tokens.input       * 1.0)
          + (tokens.output      * 5.0)
          + (tokens.cache_read  * 0.1)
          + (tokens.cache_write * 1.25)
    return ratio * input_price_per_token(model)
```

The 1.0 / 5.0 / 0.1 / 1.25 figures come from Claude’s own published per-token rates, normalised against input. That ratio holds roughly steady across the model family. What doesn’t hold is the absolute price — Opus costs several times more per token than Haiku, which is exactly what input_price_per_token(model) corrects for.

What you get, three weeks in

The first few weeks won’t tell you much. Give it time, and patterns surface: which tickets run expensive, which kinds of work quietly cost more than they should.

None of it decides whether a feature was worth building — that call still belongs to you. What changes is that you now have a real number in hand instead of a shrug, to put next to whatever value story the PM was already telling.

The question was never whether you’re spending a lot on tokens. It’s whether you’re spending them on the right things.

*Filipe Albero Pomar is an engineering manager and speaker based in London. He writes about AI adoption, engineering leadership, and software delivery. Find him at [alpomar.dev](https://alpomar.dev).*

*TechLeadConf 2026 · GenAI London 2026*
