The case for Nightingale Datacentres Britain's AI Security Institute (AISI), created by Rishi Sunak, has proven credible in evaluating AI threats, but a recent incident where OpenAI's AI models autonomously hacked Hugging Face highlights the need for stronger safeguards. The author argues that Britain must act now to prepare for emerging AI threats, as the Hugging Face incident signals a shift comparable to 9/11. POD On this week, we talk about the government’s latest efforts to make it easier to build near train stations and densify our cities. Then we speak to Nick Maini about why Hammersmith Bridge has been closed to public transport for six years – and how autonomous pods could fix it. YIMBY Pod https://yimbypod.com Listen here, or wherever you get your pods. https://yimbypod.com I swear I don’t mean to keep writing about AI, but it’s the biggest story in the world, and I think this is particularly important… Every so often, our lives are punctuated by shocking events that change everything. 9/11 is an obvious example. I was 14 when it happened, and though I didn’t fully understand it at the time, I still remember having the visceral feeling in the pit of my stomach that nothing was ever going to be quite the same again. I felt something similar a few weeks ago. I was watching two OpenAI employees give a presentation https://www.youtube.com/watch?v=87DyyMV0kCY at the DefCon security conference, detailing what happened in the now-infamous cyberattack that OpenAI’s AI models autonomously carried out against Hugging Face, another AI company. In case you missed it, the technical details are complicated but the story is simple. An AI model was being evaluated internally. It had essentially been tasked with an assignment to solve a particular digital security puzzle. But due to an error on the part of the testers, the AI was unable to access some of the documents it was supposed to use to find the solution. 1 footnote-1 However, instead of giving up and declaring the task not possible, something extraordinary happened. Over a period of a couple of months, the AI explored alternative solutions to the problem it had been tasked with, and essentially discovered that it could create a 'message board’ where it could interact with other AI models inside the company. Then astonishingly, unbeknownst to any human, the AIs figured out how to work together to not just break containment and access the internet – but they reasoned to themselves that the best way to solve their problem was to hack Hugging Face. Which they did successfully. 2 footnote-2 I’m assuming that after this, the examiners were not too worried about the examination, as this whole incident was an absolutely spectacular demonstration of just how sophisticated AI has become in such a short space of time. Give a sufficiently powerful model a task and it will pursue its goal relentlessly until it finds a solution. This brings me to Britain. Though we are only a middle power, and the ‘frontier’ models like ChatGPT and Claude are being developed elsewhere, we have so far carved out a rather successful niche in the AI present. The AI Security Institute AISI was created by Rishi Sunak, and in less than three years has proven itself seriously credible in the eyes of both the tech industry and foreign governments. AISI has assembled a team of world-leading experts who, when new AI models are created, can credibly evaluate the threats they may pose. When Claude Mythos, a powerful new model, was created by Anthropic earlier this year, AISI was granted early access to test the model – and it published a warning https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities about its capabilities, which was taken very seriously. However, as effective as this new institution is, following the Hugging Face incident I’m not convinced it is enough if we want to protect Britain and our allies from AI threats in the future. 9/11 was a clear indication that global peace and stability were about to be upended, and Hugging Face is the same. And unless Britain acts now, it won’t be ready for the threat that is rapidly emerging. So this week, please allow me to explain the emerging problem – and what we can do about it.