A curated map of where LLM tokens actually go: ~200 tools, papers and practices for token cost efficiency Quesma, a database observability company, published an open, curated list of 218 tools, papers, and practices for reducing token costs in LLM and agent workflows, covering monitoring, semantic caching, cheap local models, context engineering, and cost governance, with each entry cited and dated. The list requires cost claims to include numbers and a method, and corrections are credited. Hey all, I’ve spent the last weeks researching where tokens actually go in LLM and agent workflows: harness overhead, redundant context, cache misses, wrong model for the step. The result is an open list I maintain at Quesma: 218 entries covering monitoring, semantic caching, cheap local models, context engineering and cost governance, each with a citation and date. It’s deliberately curated, not a link dump: cost claims need numbers and a method to get in, and we say no more often than yes. If something you use is described wrong, tell me and I’ll fix it, corrections get credited. Suggestions for new entries welcome too, they go through the same review as everything else.