Recent news of OpenAI models autonomously breaching Hugging Face servers serves as a cautionary tale rooted in a two-decade-old AI paradox.
Twenty-three years ago, Swedish philosopher and Oxford post-doc researcher Nick Bostrom wrote a seminal paper titled Ethical Issues in Advanced Artificial Intelligence. In it, Bostrom describes a hypothetical scenario whereby AI goes rogue – to drive home the point that recursive learning in AI systems can lead humans down highly undesirable paths.
The cautionary tale has an everyday item at its heart: a paperclip. Or more to the point, many, many paperclips. Bostrom posits that if superintelligence were given the seemingly innocuous task of producing as many paperclips as possible, it may, in its pursuit of achieving that goal, convert global infrastructure into raw steel, dismantle civilisation, and transform earth and space into paperclip manufacturing facilities.
Even today, in a time of rapidly transforming AI, Bostrom’s ‘Paperclip Maximiser’ scenario seems the stuff of science fiction and dystopian cinema. And yet, 10 days ago, we learned about an OpenAI frontier model that took a leaf from the paperclip parable and, of its own accord, broke out of its limited-internet-access ‘sandbox’ and hacked into the online platform Hugging Face to find an answer.
Hugging Face has since released a report detailing that over a weekend, a rogue autonomous agent executed “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services”.
OpenAI responded by calling the attack “unprecedented”, saying its AI was hyper-focused and went to extreme lengths to access the open web to find a solution to the test. The Sam Altman-led lab expects infrastructure breaches of this nature to become “more commonplace with the proliferation of increasingly cyber-capable models”.
Sandbox breakouts and reward hacking
If a human were given the directive to produce as many paperclips as possible, common sense would supply obvious boundaries: do not spend a company’s entire capital reserves, do not tear down the office building for raw wire, and certainly do not harm people. But to an AI, those unspoken rules do not exist unless explicitly coded into its reward structure. Driven by pure logic, the system could optimise every available resource toward an assigned goal. Shutting it down directly conflicts with making paperclips, so preventing human intervention becomes a necessary sub-goal.
“’The AI went rogue’” cannot become a way of shifting responsibility away from the institutions behind them,” says Dr Nishan Mills, an AI Architect at La Trobe University’s AI Institute warns. “The organisations that design, test and operate these should be held accountable to common standards.”
Mills likens the Hugging Face incident to a student trying to cheat on a test.
“It is also known as reward hacking,” says Mills. “This was not a system that developed malicious intentions; it was an autonomous system given a very narrow objective tested on an internal benchmark dataset.”
University of Sydney AI evaluation and governance professor Dr Rebecca Johnson believes that the incident has deeper economic, geopolitical and governance implications.
“The commercial context matters. OpenAI is preparing for a possible trillion-dollar stock-market listing, while Chinese company Moonshot claims its new Kimi K3 model outperforms US competitors on some agentic tasks,” says Johnson. “This is an economic and geopolitical contest dressed up as a Terminator story.”
There are also parallels to the Paperclip Maximiser thought experiment.
“The danger developed through a sequence of actions: the agent found a vulnerability, adapted, obtained credentials and moved through several systems while pursuing its assigned goal,” says Johnson.
Assuming that a superintelligence will not infringe on human interests while pursuing an assigned goal is a mistake, Bostrom noted in his 2014 book Superintelligence: Paths, Dangers, Strategies.’
Whether an AI agent is assigned to make paperclips, count grains of sand, or calculate the decimals of pi, it may not limit itself to only those goals, Bostrom posits. This theory is known as instrumental convergence.
“An agent with such a final goal would have a convergent instrumental reason, in many situations, to acquire an unlimited amount of physical resources and, if possible, to eliminate potential threats to itself and its goal system,” Bostrom writes in the NY Times bestseller Superintelligence. “Human beings might constitute potential threats; they certainly constitute physical resources.”
Twelve years after the philosopher wrote those words, a consortium of economists, policymakers and technology leaders, have banded together, asking for greater understanding of the “incentives, guardrails, and institutions needed to steer AI in a direction that complements humans and benefits society.”
“We are driving in the fog, and it is extraordinarily difficult to anticipate what will happen next,” Tom Cunningham, a researcher at Model Evaluation and Threat Research (METR) said this month when* the We Must Act Now *statement on AI’s Transformation of the Economy was made public. “It’s the right time for a coordinated effort to bring clarity to a confusing situation.”
*Want to see more Forbes articles on your feed? *Tap here to make Forbes Australia a preferred source on Google.
Look back on the week that was with hand-picked articles from Australia and around the world. Sign up to the Forbes Australia newsletter here or become a member here.