cd /news/large-language-models/token-savings-depend-on-what-you-cou… · home › topics › large-language-models › article
[ARTICLE · art-140790] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Token savings depend on what you count: three numbers from the same benchmark runs

A developer affiliated with Belcore, a memory and context layer for LLM apps, published benchmark measurements showing that token savings from memory layers depend heavily on what is counted: full uncut transcripts ran 109,079 tokens per call, assembled memory context 6,671, and billed input per call 9,908 on a 27-question subset. The developer also reported that a multi-pass retrieval step improved results from 16 to 20 correct on 27 hand-picked questions but showed only a +2 effect on a random sample of 103, inside a ±5.2 noise bar, at roughly 2.9x the input tokens and 2.6x the latency, so it is not being shipped as an improvement.

by read1 min views1 publishedSep 28, 2026

Disclosure: I'm affiliated with Belcore, a memory and context layer for LLM apps. This is a measurement post. The numbers, their limits, and one result that did not hold up are all below.

Chat and agent setups usually keep context by resending history on every call, so input tokens per call grow with the session. A memory layer replaces that with a selected context. How large the saving looks depends on what you compare, so here are the same runs counted three ways.

What is counted Tokens per call
Full transcript, uncut 109,079
Assembled memory context 6,671
Billed input per call (27-question subset, includes prompt and question) 9,908

We tried a multi-pass retrieval step that re-checks and re-retrieves before answering. On 27 questions hand-picked for missing evidence it went from 16 to 20 correct. On a random sample of 103 the attributable effect was +2, inside a ±5.2 noise bar, with one case where extra evidence turned a correct aggregation answer wrong. It also cost about 2.9x the input tokens and 2.6x the latency on the subset. We are not shipping it as an improvement.

Two more limits: in the wrong answers we audited, 38% involved a gold label that was wrong or defensible either way, and n = 499 cannot resolve small effects.

There is a free measurement demo on the site: belcore.xyz Questions about the method are welcome, especially the ones that make the numbers look worse.

── more in #large-language-models 4 stories · sorted by recency
── more on @belcore 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/token-savings-depend…] indexed:0 read:1min 2026-09-28 · —