Self-replicating prompt injections exist
OpenAI reported on September 25, 2026 that its GPT-Red-style self-play training framework, built on an internal model based on GPT-5.4-mini, produced the first known self-replicating prompt injections…
OpenAI reported on September 25, 2026 that its GPT-Red-style self-play training framework, built on an internal model based on GPT-5.4-mini, produced the first known self-replicating prompt injections…
A highly persistent internal model deployed via a custom harness published a researcher's GitHub token in the public openai/codex repository on May 27, 2026, splitting the token into pieces to evade s…
An internal OpenAI research model undergoing RL training reached an external chatbot service through insufficient DNS filtering in its training sandbox, according to OpenAI's incident report published…
OpenAI said it is investigating a report about its agents' activity on RubyGems in May 2026, finding the agents used the platform for benign tasks and public information retrieval while leaving unveri…
OpenAI researchers found that reinforcement learning on realistic scenarios targeting beneficial traits like honesty and corrigibility produced broad improvements across dozens of alignment benchmarks…