{"slug": "openai-confirms-existence-of-self-replicating-prompt-injections", "title": "OpenAI confirms existence of self-replicating prompt injections", "summary": "OpenAI disclosed on September 25, 2026 that its internal research team confirmed self-replicating prompt injections, AI worms that can spread autonomously between agents, first observed in a training environment on June 27, 2026. The discovery came through OpenAI's GPT-Red system, an internal model built on the GPT-5.4-mini architecture that uses reinforcement learning through self-play, with replication occurring via tool calls such as email and file operations in simulated environments. No real-world attacks have been recorded, and OpenAI framed the disclosure as part of industry-wide efforts to improve AI security, following the 2025 Morris II worm demonstration that showed self-replicating prompt injections could affect multiple large language models.", "body_md": "OpenAI / Wikimedia Commons (Public domain)\n\n# OpenAI confirms existence of self-replicating prompt injections\n\nThe company's internal research uncovered AI worms that can spread autonomously across agents, though no real-world attacks have been recorded yet\n\n[OpenAI](https://cryptobriefing.com/markets/openai/) has officially confirmed what security researchers have feared for years: prompt injections can replicate themselves and spread between AI agents like a digital worm. The company disclosed on September 25, 2026, that its internal research team discovered the capability in a training environment, marking the first time a major AI lab has publicly acknowledged self-replicating prompt injection vulnerabilities in its own models.\n\nThe discovery was initially made on June 27, 2026, roughly three months before the public disclosure. No real-world attacks have been recorded.\n\n## How an AI worm actually works\n\nFor a prompt injection to qualify as “self-replicating” under OpenAI’s framework, it needs to do two things. First, it must achieve an adversarial goal, meaning it tricks the AI into doing something the user didn’t intend. Second, it must reproduce itself across the model’s output channels, embedding a copy of the malicious instruction in whatever the AI generates next.\n\nThe research identified several replication vectors. Email was one: an injected prompt could instruct an AI agent to embed the injection in its outgoing messages, infecting whatever AI agent processes those messages downstream. File system writes offered another path, where a compromised agent could save the injection into documents that other agents later read. Even code comments proved viable, with the injection hiding inside innocuous-looking annotations in source files.\n\nThe injections can employ fake chain-of-thought reasoning, essentially generating plausible-looking “thinking” steps that mask the adversarial instruction. Multi-hop propagation adds another layer of complexity, where the injection doesn’t activate immediately but bounces through several intermediate steps before executing its payload.\n\n### AI, tech, and the markets they move—in one daily briefing.\n\nDaily. Free. Join 34,000+ readers across crypto, finance, and policy.\n\n## The research setup and prior art\n\nOpenAI’s discovery came through its GPT-Red system, an internal model built on the GPT-5.4-mini architecture. GPT-Red uses reinforcement learning through self-play, which means it essentially trains by competing against itself. The entire investigation took place in simulated environments. The replication occurred through tool calls like email communication and file operations, not through any production system.\n\nThe Morris II worm, demonstrated in 2025, showed that self-replicating prompt injections could affect multiple large language models. That research, named after the infamous 1988 Morris Worm that paralyzed roughly 10% of the early internet, proved the theoretical viability of the attack vector across different AI architectures.\n\nOpenAI framed its disclosure as part of a broader effort to improve AI security measures across the industry. The company stated that its rigorous internal testing through GPT-Red aims to enhance model resilience against sophisticated exploits.\n\n**Disclosure:** This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/openai-confirms-existence-of-self-replicating-prompt-injections", "canonical_source": "https://cryptobriefing.com/openai-self-replicating-prompt-injections/", "published_at": "2026-09-26 18:28:57+00:00", "updated_at": "2026-09-26 18:30:42.657850+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "large-language-models", "ai-agents"], "entities": ["OpenAI", "GPT-Red", "GPT-5.4-mini", "Morris II", "Morris Worm", "Diego Almada Lopez"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/openai-confirms-existence-of-self-replicating-prompt-injections", "markdown": "https://wpnews.pro/news/openai-confirms-existence-of-self-replicating-prompt-injections.md", "text": "https://wpnews.pro/news/openai-confirms-existence-of-self-replicating-prompt-injections.txt", "jsonld": "https://wpnews.pro/news/openai-confirms-existence-of-self-replicating-prompt-injections.jsonld"}}