Has AI started improving itself?
Meta CEO Mark Zuckerberg said on 30 July 2025 that the company has 'begun to see glimpses of our AI systems improving themselves,' calling the progress slow but undeniable and saying superintelligence…
Meta CEO Mark Zuckerberg said on 30 July 2025 that the company has 'begun to see glimpses of our AI systems improving themselves,' calling the progress slow but undeniable and saying superintelligence…
OpenAI has released GPT-Red, a tool that automates prompt-injection testing against AI agents, addressing the security risks of agents that take actions in real systems. The tool helps identify poison…
OpenAI unveiled GPT-Red, its strongest automated safety red-teaming model, which successfully attacked GPT-5.1 in 84% of test cases compared to 13% for human red-teamers. The internal-only model, trai…
OpenAI introduced GPT-Red on July 15, 2026 as an internal automated red-teaming model, but a developer argues small teams can benefit from a simpler 55-line replay harness that turns prompt-injection …
OpenAI trained GPT-Red, an internal automated red-teaming model using self-play reinforcement learning, which beat human red-teamers 84% to 13% on a replicated indirect prompt injection arena and disc…
OpenAI on July 16 disclosed GPT-Red, an automated red-teaming system that uses adversarial self-play reinforcement learning to probe its own AI models for prompt injection vulnerabilities, achieving a…
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks, automating red-teaming safety evaluations typ…
OpenAI's automated red-teaming model GPT-Red outperformed human red teamers on a prompt injection test, using self-play reinforcement learning to find weaknesses in GPT models. The attacker model earn…
OpenAI has built GPT-Red, an automated red-teaming model that hunts for security flaws in its own AI systems, and will not release it due to safety concerns. In tests, GPT-Red cracked 84% of attack sc…
OpenAI introduced GPT-Red, an automated AI red-teaming system that helped strengthen GPT-5.6 against prompt injection attacks before deployment. The company reported that GPT-Red succeeded in 84% of i…
OpenAI's internal model GPT-Red outperforms human experts in AI security, finding successful attack vectors in 84% of test scenarios compared to humans' 13%, according to sources. The results are fast…
OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models, achieving successful attacks in 84 percent of test scenarios versus 13 percent for human red team…
OpenAI introduced GPT-Red on July 15th, an internal automated red-teaming model designed to find prompt injection vulnerabilities at scale and feed those attacks back into the training of production m…
OpenAI has introduced GPT-Red, an automated red teaming system that uses self-play to enhance AI safety and alignment by stress-testing AI systems against vulnerabilities such as prompt injections and…
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest ver…