{"slug": "safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents", "title": "Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents", "summary": "A new arXiv paper (2608.27141) demonstrates that safety monitors for autonomous large language model (LLM) agents fail to detect attacks whose evidence is spread across multiple iterations, because trajectory-scoped monitors re-initialize state each loop. The authors introduce LoopHarness, which maintains persistent, non-decaying safety state at the loop level, bounding unauthorized irreversible actions by B+m-1+m/δ_M, a constant independent of the horizon N. The paper provides an evaluation protocol on Agent-SafetyBench and an adaptive white-box red team.", "body_md": "# Computer Science > Cryptography and Security\n\n[Submitted on 27 Aug 2026]\n\n# Title:Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents\n\n[View PDF](/pdf/2608.27141)\n\n[HTML (experimental)](https://arxiv.org/html/2608.27141v1)\n\nAbstract:Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon $N$. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor $\\delta_M$, it bounds the expected number of unauthorized irreversible actions by $B+m-1+m/\\delta_M$, a constant in $N$, of which the $B+m-1$ term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents", "canonical_source": "https://arxiv.org/abs/2608.27141", "published_at": "2026-08-29 02:07:07+00:00", "updated_at": "2026-08-29 02:18:05.163870+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence"], "entities": ["arXiv", "LoopHarness", "Agent-SafetyBench"], "alternates": {"html": "https://wpnews.pro/news/safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents", "markdown": "https://wpnews.pro/news/safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents.md", "text": "https://wpnews.pro/news/safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents.txt", "jsonld": "https://wpnews.pro/news/safety-does-not-compose-non-decaying-loop-state-for-autonomous-llm-agents.jsonld"}}