{"slug": "token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark", "title": "Token savings depend on what you count: three numbers from the same benchmark runs", "summary": "A developer affiliated with Belcore, a memory and context layer for LLM apps, published benchmark measurements showing that token savings from memory layers depend heavily on what is counted: full uncut transcripts ran 109,079 tokens per call, assembled memory context 6,671, and billed input per call 9,908 on a 27-question subset. The developer also reported that a multi-pass retrieval step improved results from 16 to 20 correct on 27 hand-picked questions but showed only a +2 effect on a random sample of 103, inside a ±5.2 noise bar, at roughly 2.9x the input tokens and 2.6x the latency, so it is not being shipped as an improvement.", "body_md": "*Disclosure: I'm affiliated with Belcore, a memory and context layer for LLM apps. This is a measurement post. The numbers, their limits, and one result that did not hold up are all below.*\n\nChat and agent setups usually keep context by resending history on every call, so input tokens per call grow with the session. A memory layer replaces that with a selected context. How large the saving looks depends on what you compare, so here are the same runs counted three ways.\n\n| What is counted | Tokens per call | \n|---|---|\n| Full transcript, uncut | 109,079 | \n| Assembled memory context | 6,671 | \n| Billed input per call (27-question subset, includes prompt and question) | 9,908 | \n\nWe tried a multi-pass retrieval step that re-checks and re-retrieves before answering. On 27 questions hand-picked for missing evidence it went from 16 to 20 correct. On a random sample of 103 the attributable effect was +2, inside a ±5.2 noise bar, with one case where extra evidence turned a correct aggregation answer wrong. It also cost about 2.9x the input tokens and 2.6x the latency on the subset. We are not shipping it as an improvement.\n\nTwo more limits: in the wrong answers we audited, 38% involved a gold label that was wrong or defensible either way, and n = 499 cannot resolve small effects.\n\nThere is a free measurement demo on the site: [belcore.xyz](https://belcore.xyz/?utm_source=devto&utm_campaign=token-counting)\n\nQuestions about the method are welcome, especially the ones that make the numbers look worse.", "url": "https://wpnews.pro/news/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark", "canonical_source": "https://dev.to/member_b8352c00/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark-runs-4fje", "published_at": "2026-09-28 04:48:48+00:00", "updated_at": "2026-09-28 05:17:58.232680+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "mlops", "ai-infrastructure"], "entities": ["Belcore"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark", "markdown": "https://wpnews.pro/news/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark.md", "text": "https://wpnews.pro/news/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark.txt", "jsonld": "https://wpnews.pro/news/token-savings-depend-on-what-you-count-three-numbers-from-the-same-benchmark.jsonld"}}