{"slug": "when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse", "title": "When Fancy Eviction Fails: Rethinking Cache Replacement for LLM Prefix Reuse", "summary": "A September 24, 2026 arXiv paper studying production traces from two companies found that 14 eviction algorithms tested for LLM prefix caching deliver little benefit over LRU despite a large gap to Belady, because prefix reuse is dominated by the regular pacing of active sessions. The authors attribute the result to recency being unusually predictive under agentic workloads and introduce the compute-savings ratio plus two offline oracles to quantify heavy-tailed session footprints and variable miss costs. They recommend retaining recency as the foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity, and plan to release the traces and simulator.", "body_md": "# Computer Science > Distributed, Parallel, and Cluster Computing\n\n  [Submitted on 24 Sep 2026]\n\n# Title:When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse\n\n[View PDF](https://arxiv.org/pdf/2609.28870)\n\n[HTML (experimental)](https://arxiv.org/html/2609.28870v1)\n\nAbstract:Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.\n    \n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse", "canonical_source": "https://arxiv.org/abs/2609.28870", "published_at": "2026-10-01 21:44:44+00:00", "updated_at": "2026-10-01 22:15:48.424319+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-research"], "entities": ["arXiv", "Belady", "LRU", "HBM"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse", "markdown": "https://wpnews.pro/news/when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse.md", "text": "https://wpnews.pro/news/when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse.txt", "jsonld": "https://wpnews.pro/news/when-fancy-eviction-fails-rethinking-cache-replacement-for-llm-prefix-reuse.jsonld"}}