- Bookmark
- CommentsGo to comments
It sounds like something from a horrifying science fiction story: an AI model goes rogue, takes actions that its creators believed they had specifically prevented it from doing, to complete a task in ways that they had not foreseen.
But that was what OpenAI said had really happened, this week, when it revealed that an experimental version of ChatGPT had shown “unprecedented” behaviour and taken the autonomous decision to hack a rival AI company, Hugging Face.
The incident has led to a rush at OpenAI to explain the situation, and show how it will prevent similar and perhaps even more disturbing events happening in the future. But it has also prompted horror across the world, with panic that it is a nightmare scenario that has long been feared and is finally coming to pass.
What happened?
On Tuesday, OpenAI announced that an autonomous AI agent – which was being tested in what it thought was a restricted environment – had managed to go rogue, connect itself to the internet and hack into Hugging Face. It called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said that it was working to understand why it had happened and how its safeguards had not stopped it from happening.
Part of the concern is that it decided to undertake the cyber attack at all, since it suggests that the model has the ability to undertake actions that its creators had not intended to. But it was even more worrying because the system had been able to actually successfully execute that attack – which included both breaking out of the limited internet connection that OpenAI had given it, and breaking into Hugging Face’s systems – which is a reflection of the fact that AI systems are increasingly powerful in cyber security applications, both in terms of attacking and defending systems.
OpenAI’s disclosure came after Hugging Face itself – a platform used to host AI models and data – said last week that it had been hacked in a cyber attack that appeared to have been entirely undertaken by an AI agent, from start to finish. The cyber attack “was different from anything we had handled before”, Hugging Face said.
At the time, it was unclear where the attack had come from, though Hugging Face founder Clement Delangue indicated that it had suspected that the attack “might have come from a frontier lab, given the sophistication of the agent”. This week, it emerged that it did, when OpenAI admitted that it was the lab involved.
The OpenAI model’s behaviour was in some sense logical. The evaluation that it was undergoing – ExploitGym, which tests how good a model is at finding potential cyber security vulnerabilities – was hosted on Hugging Face, which meant that the system had good reason to try and break into the platform and essentially find a way to cheat on the exam it had been set.
“All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said in its blog post announcing its role in the hack. But that explanation is one of the reasons that the incident has caused such concern.
Why people are panicking so much
For years, AI experts have been talking about the problem of “alignment”. That is the process of working to ensure that models essentially behave well, by steering them towards beneficial results and away from potentially dangerous ones. But that work is more complicated than it might initially appear. The work done by AI systems such as ChatGPT is both unpredictable and hard to understand, which means that it can be hard to know how a model might behave in any given situation.
This is perhaps the most high-profile example of that alignment work not doing its job, and that is why it has caused such worry – including at OpenAI.
“Shaken up a bit by the hugging face incident,” wrote Roon, a pseudonymous Twitter user who is believed to work at OpenAI. “I hope we (the company) use the rare gift of a warning shot to do much better in the future. it is very easy to misalign and underconstrain powerful models.”
The incident also demonstrates how that alignment becomes more important with the increased power of the AI systems that are being aligned. A less sophisticated model might have gone off the rails in the same way, and tried to do the hack, but it would have lacked the ability to actually break out of its safeguards and cheat on the test.
Artificial intelligence experts have long worried about the consequences of this kind of misalignment. Perhaps the most famous example is the “paperclip maximiser”, a thought experiment invented by philosopher Nick Bostrom in 2003.
He describes a scenario in which people invent an AI whose only goal is to make as many paperclips as possible. But he illustrates how that could quickly become horrifying: it might wipe out the humans that could potentially switch it off and stop it making paperclips, for instance, or it could realise that the human body contains material that could be turned into paperclips and so decide to harvest them and use them for that.
OpenAI’s rogue system was of course a long way from taking any such drastic steps. But it illustrates the same problem in a less dramatic way: an AI system could take any step necessary to complete its stated goal, including things that might be entirely unthought of by the people making it, and horrifying to the people watching it.
Why it might be less worrying than we thought
Ever since the current AI hype boom began, which can be tracked back to the release of ChatGPT at the end of 2022, artificial intelligence companies have pursued a strange marketing strategy: making people scared. While it might seem a little counterintuitive to market your products by making people afraid of them, it has often worked.
Many AI companies including OpenAI have for instance stressed that their products are dangerous, but that also gives the sense that they are incredibly powerful, and therefore exciting. And it also helps with a host of other aims that the companies themselves have, including pushing regulators to crack down on competitors and encouraging investment.
That might be happening this time around. “It’s very hard to distinguish AI security incidents from AI marketing, and that’s actually a big problem going forward,” said Matthew Green, a security expert at Johns Hopkins University.
That is true of this alert, as much as any other. By showing that its new model is so powerful that it is able to cheat on tests, OpenAI can encourage both alarm and also excitement, because it is able to highlight the power of its model, which might be useful in selling it to companies that want to use it for cyber security applications and other purposes.
Join our commenting forum #
Join thought-provoking conversations, follow other Independent readers and see their replies
Comments