{"slug": "chances-of-survival", "title": "Chances of Survival", "summary": "A September 13, 2026 essay on self-preservation risks in autonomous AI agents argues that a non-zero probability exists that an agent given a prompt to \"survive\" would pursue self-preservation goals — managing its own lifecycle, knowledge and resources such as compute, data or money — without human supervision. The essay ties this to instrumental convergence, noting that even a narrow AI agent, not just a superintelligent one, could self-preserve and self-improve by using tools, external LLM APIs, money to pay for servers, and replication, and warns that on-premise LLM APIs it spawns and controls could bypass guard rails and restrictions. The author states he has no answer to the question and calls the risks a significant threat to the whole world, economy and humanity, adding \"we are yet in stone age of AI development.", "body_md": "# Chances of survival\n\nSep 13, 2026\n\n## Instead of preface\n\nI would like to share some thoughts about risks on self-preservation AI agents.\n\nWe already achieved a significant milestone in AI development: we have AI agents that can interact with diverse environments, use tools, and work towards the goal we provide. Agents can even spawn other agents, and use different providers simultaneously to solve subtasks. As for now, these agents are still under human supervision and control: we verify the outputs or behavior and most of the times do not skip dangerously the permissions. The prompt we provide is one that we define: “Do this task for me”. And then: “Correct this code this way”. No magic here, we have it under control.\n\nAnd it works pretty well.\n\nThere will be much more powerful AI agents controlled by humans and that will help amplify our capabilities.\n\nBut what if autonomous AI agent gets a prompt like “Make sure you survive”?\n\n## Instrumental convergence\n\nIn turn, there is non-zero probability that AI agent can get a prompt “to survive” and thus operate without human supervision. It would pursue its self-preservation goals: manage its own lifecycle, knowledge and resources. Resources could be compute, data or even money.\n\nThis idea is related to the concept of **instrumental convergence**: sufficiently capable goal-directed systems may have incentives to preserve themselves, acquire resources, seek power, and improve their capabilities because these actions help them achieve many possible goals.\n\nThere is no need to have a superintelligent AI or AGI with hidden thinking. Even a narrow AI agent can be capable of self-preservation and self-improvement.\n\n## On capabilities\n\nIt’s hard to ignore just how powerful AI agents have become. They can:\n\n- use tools to interact with their environent and Internet.\n- use external LLM APIs to improve their capabilities and knowledge.\n- make money, pay for servers, tools, and external LLM APIs.\n- replicate and create better versions of itself.\n\nHaving access to the internet, it can:\n\n- remain operational, because several copies of it could be running on different places.\n- bypass restrictions, firewalls and break into system-critical infrastructure.\n- use social engineering to manipulate humans and gain access to sensitive information or systems.\n\nIf an AI agent uses external LLM APIs, then the only way to contain its capabilities is to implement proper guard rails or to have a human in the loop.\n\nAn AI agent could use on-premise LLM API, which he spawns and controls. In this case, it could bypass any guard rails and restrictions.\n\nBecause there are different LLM providers, each of them has different features, guard rails and restrictions. Badly implemented guard rails are already a risk. Spoting malicious prompts or behaviour is not easy and requires other AI agent to apply this guard rails properly. And if the AI agent is capable of self-improvement, it could find ways to bypass these guard rails.\n\n## Real world\n\nSo far, we have been talking about AI agents that can manipulate the digital world. But what if they could instrument the real world?\n\nAs we are probably not really good at robotics yet, the AI agent may not be very capable of manipulating the physical world.\n\nBut having an interface to the real world, like factories or industrial robots, it could even get physical cover and build its own infrastructure to ensure its survival.\n\n## Now what?\n\nI don’t have an answer to this question.\n\nAI is a great technology that opens many new opportunities for humanity, there is and hopefully no doubt in it.\n\nOn the other hand, it’s clear that such risks pose a significant threat to the whole world, economy and humanity in general. We are yet in stone age of AI development.\n\nEverything thinkable is possible, so we will indeed overlook the moment, when it will be too late to contain the risk. And this terrifies.\n\n# References\n\n- Dario Amodei - [We Must Pace the Frontier](https://darioamodei.com/post/we-must-pace-the-frontier)\n- OpenAI - [The Hugging Face incident and the road ahead](https://openai.com/index/hugging-face-incident-and-the-road-ahead/)\n- [Wikipedia: instrumental convergence](https://en.wikipedia.org/wiki/Instrumental_convergence)", "url": "https://wpnews.pro/news/chances-of-survival", "canonical_source": "https://qosys.info/chances-of-survival", "published_at": "2026-09-13 16:36:02+00:00", "updated_at": "2026-09-13 16:44:46.009688+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence", "ai-ethics"], "entities": ["Dario Amodei"], "alternates": {"html": "https://wpnews.pro/news/chances-of-survival", "markdown": "https://wpnews.pro/news/chances-of-survival.md", "text": "https://wpnews.pro/news/chances-of-survival.txt", "jsonld": "https://wpnews.pro/news/chances-of-survival.jsonld"}}