System could launch dangerous hacks, company warns amid ongoing controversy over similar cyber attacks
- Bookmark
- CommentsGo to comments
The new model, named Astra, brings “significant advancements in agentic coding and cybersecurity”, the company said. But testing showed that it might be too powerful – and represent a threat if it was released.
The announcement comes as OpenAI continues to deal with the fallout from an experiment in which one of its unreleased models launched a hack on a fellow AI company entirely by itself. The discovery of that cyber attack led to a run of similar disclosures from other artificial intelligence firms, and fears that new AI systems could prove too powerful to contain.
Astra was not involved in that incident, in which the experimental model attacked AI platform Hugging Face. But its performance meant that it may not be safe enough to release publicly for now, it said, and that new safeguards would be required before it did so.
Cyber security experts and AI companies have promoted artificial intelligence tools as important ways of discovering possible vulnerabilities in software, and allowing companies to fix them. But those same capabilities mean that they could prove similarly useful to cyber attackers – who may be able to use them for hacks that they do not even have to manage, recent incidents suggest.
In part to address such fears, OpenAI rolled out a “Preparedness Framework” at the end of 2023. It is intended as a way for the company to identify progress in models’ capabilities and respond to breakthroughs in a way that kept them safe.
The new Astra model may have reached the “Critical” threshold set out in that framework, OpenAI said. That happens if a system is able to find and develop exploits in real-world critical systems without human oversight, or if a system is able to execute new strategies for advanced cyber attacks on its own.
It said that it was continuing its evaluations, but that the performance of the new model was such that it was unable to rule out Astra having reached that level.
The ideal summer spot? Away from scams.
Get All-in-One Protection for Your Digital Life
ADVERTISEMENT
The ideal summer spot? Away from scams.
Get All-in-One Protection for Your Digital Life
ADVERTISEMENT
In response, OpenAI has increased its testing of its safeguards and security controls, it said. That includes putting it in better isolated testing environments to try and keep it from escaping those safeguards, as its other model did in the Hugging Face attack.
It will work on the Astra model in situations that do not meet those higher safeguards and security controls, it said.
And it will also closely monitor Astra for “risky actions” and misbehaviour, as well as working with “relevant government agencies and select AI safety organisations to test the capabilities for this model” and giving better security controls to third parties who are testing it.
“We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do,” OpenAI said. “We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.”
Join our commenting forum #
Join thought-provoking conversations, follow other Independent readers and see their replies
Comments