{"slug": "kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning", "title": "KV-streams for Efficient Compaction in Agentic Reinforcement Learning", "summary": "Researchers Emiliano Penaloza and co-authors submitted KV-streams to arXiv on 28 Sep 2026, a plug-and-play context-compaction strategy for agentic reinforcement learning that streams the KV cache forward instead of flushing it after each compaction. KV-streams enables three different compaction strategies and achieves a 2.6 to 5x wall-clock training speedup, with the streamed KV cache acting as a recurrent state that carries forward information dropped from context; the authors report that RL alone is sufficient for this behavior to emerge, contrary to prior work.", "body_md": "# Computer Science > Machine Learning\n\n  [Submitted on 28 Sep 2026]\n\n# Title:KV-streams for Efficient Compaction in Agentic Reinforcement Learning\n\n[View PDF](http://arxiv.org/pdf/2609.35750v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.35750v1)\n\nAbstract:Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.\n    \n\n## Submission history\n\nFrom: Emiliano Penaloza [\n[view email](http://arxiv.org/show-email/bb631b20/2609.35750)]\n\n**[v1]** Mon, 28 Sep 2026 17:57:42 UTC (17,726 KB)\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\nIArxiv Recommender\n\n*(*[What is IArxiv?](https://iarxiv.org/about))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning", "canonical_source": "http://arxiv.org/abs/2609.35750v1", "published_at": "2026-09-29 17:04:28+00:00", "updated_at": "2026-09-29 17:20:43.839320+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research", "ai-agents", "ai-infrastructure"], "entities": ["KV-streams", "Emiliano Penaloza", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning", "markdown": "https://wpnews.pro/news/kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning.md", "text": "https://wpnews.pro/news/kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/kv-streams-for-efficient-compaction-in-agentic-reinforcement-learning.jsonld"}}