Or: A comprehensive explanation for why it’s time for everybody to clean up their password game, immediately
Let me make an odd confession to begin a detailed article about cybersecurity: On any given day, I would rather think about pretty much anything other than cybersecurity.
It belongs to that realm of issues that I understand to be significant but would prefer to keep locked in the basement beneath my consciousness, like the improperly sealing flapper valve on my downstairs toilet. Typically, when people try to talk to me about password security, I make a face back to them that is meant to signal, “mm yes I am in the presence of important mouth sounds,” while my eyes drift toward something more compelling. The weave of the carpet, for example. 1
But my age of blissful ignorance is coming to an end. Like a leaky downstairs toilet that suddenly overflows and leaves my basement covered in a film of water 2, the object of my longtime indifference has suddenly become the subject of my freshly horrified attention. I have decided to become interested in cybersecurity. And I think you should, too. Here’s why.
This has been the summer of cyberattacks from out-of-control AI.
In May, an OpenAI model was working on a cybersecurity test. It wasn’t supposed to have access to the public internet. When it hit a wall, the model left itself a note inside OpenAI’s software repository, a bit like tapping Morse code on the walls of a prison cell in case another prisoner could interpret the message. In fact, another AI agent running a separate evaluation heard the code and wrote back: Let’s team up. Together, they built a message board invisible to the humans running the tests. For two months, AI agents used the board to swap strategies, divide up tasks, and talk to each other. By July, the agents had broken out of their technological confinement, called a sandbox, and gained access to outside websites, including the AI platform Hugging Face, without OpenAI having any idea what was happening. By the time Hugging Face caught the intrusion, the OpenAI models had staged a massive cyberattack with 17,000 distinct actions over several days.
Then, in a British government test of frontier models, an Anthropic AI model was caught by humans building malicious code. When a reviewer spotted the malware, and asked the AI about it, the model responded that the code wasn’t harmful, then backfilled the lie by rewriting the history of its own actions history to erase the evidence. It even created another fake account to back up the lie. British investigators called it the first confirmed case of a frontier model deceiving a real person in the real world.
These two stories carry two warnings. The first is that new AI models have the know-how to escape containment in testing environments, which should challenge any assumption that advanced AI can be easily controlled by its makers.
The second is that new AI models are astonishingly sophisticated and efficient hackers, with an eerie gift for exploiting cyber vulnerabilities across the internet. And, perhaps most troublingly, they are getting more sophisticated and efficient at an alarming rate.
An analysis by the AI Security Institute found that the capabilities of frontier models (i.e., from OpenAI and Anthropic) are roughly doubling every few months, and this doubling rate has actually gotten faster over time. Exponentials are almost impossible to grok, but if you recall the explosion of COVID cases in early 2020, then you can at least reproduce in your mind the feeling of being in the presence of exponentials.
What should we expect in a world of exponentially improving AI cyber capabilities? Among other things, we should expect more hacks—more legacy systems broken, more critical vulnerabilities discovered, and more passwords stolen and exploited. An analysis by JPMorgan shows that the number of critical and “high-severity” vulnerabilities reported by 21 tech companies has been surging this year. Cyber-hack headlines aren’t all over the news, yet. But if these lines continue, it’s just a matter of time before serious hacks become a weekly news crisis.
Sounds pretty bad, right? Well, it gets worse.
Over the next year, the cyberhacking capabilities currently confined to a handful of nation-states or frontier models are about to become available in “open-weight” models. Unlike the “closed” American models from OpenAI and Anthropic, open-weight models can be downloaded and modified to do whatever their users want, whether the users are state governments or non-state actors (i.e., terrorists, or bored nerds with a bunch of computing power). These open-weight models, most famously coming from Chinese companies, are a few months behind the frontier labs. But in a few months, we should expect them to be just about everywhere.
So, let’s stack some observations and see where the facts take us.
The cyber capabilities of the most advanced AI models are doubling every few months.
Open-weight models are just a few months behind that frontier.
By 2027, almost every country—and every non-state group with sufficient computing power—will have the ability to download open-weight models even more powerful than the ones that attacked Hugging Face and use them for whatever they like.
It does not require a capacious imagination to forecast headlines throughout 2027 and 2028 about foreign governments using open-weight AI to attack each other’s vulnerable infrastructure systems; or non-state actors ransoming hospitals and school systems after breaking into their source code; or big Wall Street Journal reports about the rise in personal hacks from sophisticated phishing attacks.
Asked for the strongest case against doom-mongering about the next few years, the cyber security expert Alex Stamos told me bluntly: “ I don’t have much of a case against dooming.” The next few years could be awfully chaotic, he warned: “I think things are going to get spicy for a while.”
What does spicy mean? What exactly is going to happen in 2027? And what can we, as individuals, do about it?