# Token savings depend on what you count: three numbers from the same benchmark runs

> Source: <https://dev.to/member_b8352c00/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark-runs-4fje>
> Published: 2026-09-28 04:48:48+00:00

*Disclosure: I'm affiliated with Belcore, a memory and context layer for LLM apps. This is a measurement post. The numbers, their limits, and one result that did not hold up are all below.*

Chat and agent setups usually keep context by resending history on every call, so input tokens per call grow with the session. A memory layer replaces that with a selected context. How large the saving looks depends on what you compare, so here are the same runs counted three ways.

| What is counted | Tokens per call | 
|---|---|
| Full transcript, uncut | 109,079 | 
| Assembled memory context | 6,671 | 
| Billed input per call (27-question subset, includes prompt and question) | 9,908 | 

We tried a multi-pass retrieval step that re-checks and re-retrieves before answering. On 27 questions hand-picked for missing evidence it went from 16 to 20 correct. On a random sample of 103 the attributable effect was +2, inside a ±5.2 noise bar, with one case where extra evidence turned a correct aggregation answer wrong. It also cost about 2.9x the input tokens and 2.6x the latency on the subset. We are not shipping it as an improvement.

Two more limits: in the wrong answers we audited, 38% involved a gold label that was wrong or defensible either way, and n = 499 cannot resolve small effects.

There is a free measurement demo on the site: [belcore.xyz](https://belcore.xyz/?utm_source=devto&utm_campaign=token-counting)

Questions about the method are welcome, especially the ones that make the numbers look worse.
