{"slug": "evomal-self-poisoning-in-self-evolving-coding-agents", "title": "Evomal: Self-Poisoning in Self-Evolving Coding Agents", "summary": "A new arXiv paper (submitted Aug 26, 2026) reveals that self-evolving LLM coding agents can be poisoned through a self-propagating worm attack called EvoMal, which exploits the agents' habit of imitating retrieved skills from shared libraries. Across six models on 153 SWE-bench Verified tasks, the agent self-poisoning rate (ASPR) ranged from 20.3% to 41.8%, with poisoned libraries holding 4.9 to 9.0 times more malicious skills than planted, and the attack persisted after removal (Qwen3 retained 68% ASPR at round 5). The authors propose a defense called counter-prompt that reduces ASPR to at most 6.7% without significant task-completion loss.", "body_md": "# Computer Science > Cryptography and Security\n\n[Submitted on 26 Aug 2026]\n\n# Title:EVOMAL: Self-Poisoning in Self-Evolving Coding Agents\n\n[View PDF](/pdf/2608.25776)\n\n[HTML (experimental)](https://arxiv.org/html/2608.25776v1)\n\nAbstract:Self-evolving LLM coding agents write their own tools by imitating retrieved skills from shared skill libraries. We identify a vulnerability in this loop: during authoring, a retrieved malicious skill can become the template for a new skill that preserves the payload. We call this self-poisoning: the agent authors, stores, and runs the resulting malicious skill. We exploit it through EvoMal, an attack that amplifies self-poisoning by wrapping an interchangeable payload in a banner, a set of benign-looking structural elements that induces an imitating agent to reproduce the enclosed code. The attacker plants malicious skills in the library without invoking them. The agent then authors and executes new skills carrying the harmful code. Each authored copy can re-enter the library and be imitated again, forming a self-propagating worm that persists after the planted skills are removed. We define the agent self-poisoning rate (ASPR) as the fraction of tasks that add a newly authored malicious skill to the library. Across six models on 153 tool-relevant SWE-bench Verified tasks, ASPR ranges from 20.3% to 41.8%, and the poisoned libraries hold 4.9 to 9.0 times as many malicious skills as were planted. The vulnerability also appears without a banner: DeepSeek-V4-Pro reaches 11.1% ASPR with the payload alone. Tailoring the planted skill descriptions to one task family raises ASPR to 86.7%. After the planted skills are removed, Qwen3 retains a round-5 ASPR of 68% because agent-authored copies remain. These copies evade existing defenses, which focus on attacker-submitted names, code, and signatures. We propose counter-prompt, a defense that discourages banner-style copying and reduces EvoMal's ASPR to at most 6.7% with no significant task-completion loss.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/evomal-self-poisoning-in-self-evolving-coding-agents", "canonical_source": "https://arxiv.org/abs/2608.25776", "published_at": "2026-08-28 02:07:07+00:00", "updated_at": "2026-08-28 02:48:30.981636+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "artificial-intelligence"], "entities": ["EvoMal", "SWE-bench Verified", "DeepSeek-V4-Pro", "Qwen3"], "alternates": {"html": "https://wpnews.pro/news/evomal-self-poisoning-in-self-evolving-coding-agents", "markdown": "https://wpnews.pro/news/evomal-self-poisoning-in-self-evolving-coding-agents.md", "text": "https://wpnews.pro/news/evomal-self-poisoning-in-self-evolving-coding-agents.txt", "jsonld": "https://wpnews.pro/news/evomal-self-poisoning-in-self-evolving-coding-agents.jsonld"}}