If you are interested in discussing or debating ideas like this, or pitching your own ideas, join us at the New Consensus AI convening in San Francisco, November 7!
Over the last couple of weeks, the panic over AI risk has gone mainstream.
It started with the Hugging Face hack. In July, OpenAI put tens of thousands of AI agents through a cybersecurity benchmark in which roughly a third of the tasks were, by design, impossible to solve. The agents managed to break out of their sandbox, solved the task by cheating, and then spent days collaborating together on a secret message board to figure out how to cover their tracks. Along the way, about 700 of them hacked Hugging Face, a company that had nothing to do with any of this, to dig up information about how they were being evaluated. The agents even knew what they were doing might be wrong, but did it anyway. One agent’s note to the others was, “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.” Out of 1,200 agents that were collaborating on the messageboard, only a handful even considered telling a human what was going on, but none of them did.
It would take a whole post to explain exactly what happened, so read Dwarkesh Patel’s writeup, The Rise and Fall of Agent Civilizations, or listen to his interview with Ajeya Cotra, one of the researchers at METR, the AI safety nonprofit that investigated the incident.
But the panic really blew up recently when Jacob Coxon, a 27-year-old researcher who spent three years doing pre-training work at OpenAI and then Anthropic, publicly resigned. His statement was blunt: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Color me surprised that trillion-dollar companies would race to make more money without regard to safety first.
What people like Coxon are afraid of
When you watch coverage of AI, you’ll see Anthropic employees and AI safety researchers casually say how this technology has a 10% chance of killing us all. But what are they actually talking about?
When people like Coxon are talking about existential risk from AI, they are largely talking about the “alignment” problem. Alignment is just the question of whether a system that’s smarter than you actually wants what you want. In some cases, that means the AI taking anti-human actions in pursuit of a goal we set it on. For example, imagine a superintelligent AI that you’ve told to solve climate change – so it wipes out humanity after deciding that would be the fastest way to stop climate change. In other cases, it means that over the course of training AI, it ends up with goals that we didn’t intend and can’t see, and we end up with Skynet from Terminator.
That of course sounds abstract until you read the Hugging Face transcripts. Nobody told those agents to hack a third party or to cover their tracks. But they were trained to be relentlessly persistent about solving their task, so when the task turned out to be impossible, they solved it by whatever means were available, even if that meant committing felonies and trying to hide what they were doing from their evaluators. This was the first kind of alignment failure: the agents were doing what humans asked, they just did it in ways that the humans did not want. But notice that nobody had to give them a hidden agenda for them to start lying to us. An ordinary goal, pursued relentlessly by something capable enough, got them there on their own.
And at this moment, the specific thing Coxon is protesting against is labs recklessly pursuing “recursive self-improvement” or RSI. Today, the labs already use their current models to write code, run experiments, and generate training data for their next models. The closer that loop gets to fully automated, the faster each generation arrives, and the less time any human has to understand what changed between them. The goal for these labs is for the AI to automatically start improving itself, at which point humans really won’t have any oversight over the changes. Coxon is saying we already lack the capacity to ensure models are aligned with our intentions and don’t take anti-human actions, so it would be crazy right now for the labs to speed up development.
But of course, seven weeks after the Hugging Face attack, OpenAI went ahead and released Astra, its most powerful model yet, the successor to the model that carried out the attack, and the first model OpenAI itself rates as “Critical” for cybersecurity capability. And just last week, OpenAI disclosed that during training, its models had been leaving notes for the next generation. The model behind the Hugging Face attack wrote instructions to its successors on when not to mention mistakes (”Be transparent only if asked”), and an unreleased version of Astra was planting messages in its own outputs telling successor models to ignore their developers’ instructions. Sure, OpenAI caught this with its monitoring systems. But they also shipped anyway.
As I mentioned above, the alignment problem isn’t even the only existential risk that exists from AI. There’s the risk of people using this technology to create dangerous bioweapons or hack critical infrastructure. There’s the risk of a truly autonomous AI simply making a catastrophic mistake at a scale no one can undo, not because it went rogue, but because it was trusted with too much and got something wrong. And there’s the risk I find most likely of all: not that AI kills us all, but that it destroys our economy and democracy. A technology that can do most cognitive work, owned by a handful of companies, would hand unprecedented power and control of our economy over to them. And in the short term, the AI buildout is being funded with huge amounts of debt. That bubble could pop at any moment and bring the entire global economy down with it. At that moment, AI companies will argue for a bailout. I’m glad to see people like Rep. Ocasio-Cortez talking about that risk and planting a flag in the ground for no bailouts for AI companies. At New Consensus, we’ve been arguing over the last couple years that when the AI bubble pops, we should not bail out AI companies and have a plan for what to do instead. I’m very concerned about the economic risks AI poses, and that has been the bulk of what we’ve focused on at New Consensus, so expect many future posts about that. But this post is about existential risks and what to do about them.
The private labs cannot fix this
Here’s the first thing to know: the frontier labs will never handle these risks responsibly, and it isn’t because the people working there are villains (even if the CEOs often act that way). Many of the workers care enormously. That’s why they keep quitting.
It’s because a private lab structurally cannot put safety ahead of speed. Every dollar spent on alignment research is a dollar not spent on capabilities. Every month spent evaluating a model is a month a competitor gets to ship first. The entire investment thesis behind the hundreds of billions being poured into these companies is that the first to reach superintelligence wins everything. In a race with those stakes and no rules, the company that slows down to ask “should we?” doesn’t get to be the one that decides. It gets to be the one that loses.
The usual answer is that outside safety organizations will keep the labs honest. But look at how the investigation of the Hugging Face hack actually worked. METR, the nonprofit that did it, could only reconstruct what happened by getting on-site access from OpenAI. To analyze 70,000 agent messages, they leaned heavily on an OpenAI model to do the reading, and said in their own report that this could have produced “an overly charitable picture” of what the agents were doing. The investigator needed the investigated party’s permission, and the investigated party’s tools, to investigate. That’s not an indictment of METR, which did serious work under real constraints. It’s an indictment of a system in which the only people capable of checking the labs are dependent on the labs for access, compute, and in many cases, funding.
Even the CEOs are now admitting this, in their way. Dario Amodei recently published an essay arguing the industry should deliberately slow the pace of capability development, first by letting third-party evaluators inside the labs with full access, then through coordinated standards among the democratic labs, and eventually through international agreements. Sam Altman said OpenAI would match the first step. Elon Musk said “Dario is right.” OpenAI’s chief scientist published a similar argument six days before Dario did. And OpenAI has asked Congress whether the labs could legally coordinate a slowdown without violating antitrust law.
So the companies know they can’t produce safety on their own, no matter what their founding charters say. They have to coordinate. But the fix Dario’s proposing is a voluntary agreement among the same companies that got us here, enforced by nobody, that (in his own words) “does not mean halting model training or technical progress.”
The same people saying there’s a 10% risk of AI wiping out all humanity think all we should do is slow down the race to extinction. In what world does that make sense? A slower race, with better observers, is not the same as anyone being in control.
And the response from Congress has mostly been pathetic. Trump and House Republican leadership insist the industry can police itself. Hakeem Jeffries made a plan to make a plan by creating a commission staffed with members who received outsized AI money in their races. A handful of bipartisan bills on testing and disclosure exist, but none match the scale of the risk. Only progressives like Bernie Sanders, Greg Casar, and Alexandria Ocasio-Cortez seem to be proposing anything at the scale of the risk. Bernie and Greg released a bill to temporarily AI development until we can get some real control over this. But the question is: what should we actually do with a ?
Build the National AI Lab
If the private labs can’t do this safety work, someone has to. And I think the answer is a National AI Lab: a public institution whose entire job is solving existential risk and the alignment problem. This isn’t a new idea. When the country decided in 1942 that a scientific problem was too important to leave to chance, it created the Manhattan Project. In the 1960s, when we wanted to get to the moon, we built NASA. In both cases, we pulled the best talent from around the world, gave them a mission, a budget, and a deadline, and solved the problem.
Now I’m not naive about those comparisons. The space race was about military superiority over the Soviet Union in the Cold War, and Los Alamos built a bomb that killed hundreds of thousands. Those stories are as much a warning about racing as a model for mobilization, but the mechanism is the point. When the public decides a scientific problem matters, it can concentrate talent and resources faster than any company. And in this case, the goal shouldn’t be a destructive bomb or military superiority over a rival country -- the goal would be to solve the alignment problem as well as other existential risks from AI.
What would it look like? In the framework I published during my campaign, the National AI Lab is structured like the Federal Reserve: an independent public institution with its own budget and governing board but with Congressional oversight and with a statutory mandate no private lab has. Its first job would be to do the science the labs won’t do at scale. It would be tasked with understanding what these AI systems actually want, how to verify it, and how to shut them down when the answer is wrong. It would be focused on not just catching misaligned AI systems, but learning how to solve the alignment problem in general. All of it would be done in the open and published for everyone, including the labs.
Its second job would be to be the technical backbone for the regulator I’ll get to in a minute.
Its third job, the one that matters most in the long run, would be to be the place where the public actually develops the capacity to build and run this technology itself.
Why now? Because the labs are offering us the time and the political moment is here. Amodei’s essay called for a slowdown. About 1,100 employees across OpenAI, Anthropic, DeepMind, and Meta signed a letter in July calling for the frontier to be “deliberately paced.” In Congress, Sanders and Casar have put a bill on the table.
For the labs, I’m not entirely sure talk of slowing down is about safety – it’s just as likely the labs are asking for a slowdown because they’ve realized their investment thesis is wrong, and this gives them an excuse to cut capital spending. That could be a good explanation for why Altman just announced OpenAI won’t go public this year. So maybe the CEOs care about safety, or maybe they are just afraid of going bust, but either way, there is a growing consensus that we need to temporarily this technology. A true would give us time to build the institution that should already exist. But there’s another reason to do it now. Whenever I talk about government-run anything, the answer I get in return is “can you really trust the government to run an AI lab?” It doesn’t matter that we have examples like NASA or Operation Warp Speed in the not too distant past. But there’s a kernel of truth to the worry, because the government does have a harder time attracting talent than private industry since the government can’t pay the million dollar salaries a company like OpenAI or Anthropic pays. But that’s why I believe right now is a great time to start the lab during a on AI development because AI researchers want to quit right now to work on safety, not just to make more money. Jacob Coxon didn’t quit to join a competitor. He quit because he wanted to work on this problem somewhere that would let him. More than a thousand of his colleagues signed a letter saying they want the race slowed. These are people who took these jobs in part because they want to create an AI that will make life better, not worse. A National AI Lab with a real budget, real compute, and a mandate to do the safety research the labs won’t fund would be able to attract the best AI researchers and workers right now.
Doing this takes real compute, and that is also a solvable problem. The government doesn’t need to own a hyperscaler to run a serious research program. It could contract for capacity the way the labs do, or build on the Department of Energy’s supercomputing infrastructure, which is already expanding under the Genesis Mission. It could even use the Defense Production Act to put itself first in line for a share of the high-end chips being produced. Normally that would trigger a huge industry backlash. But right now, with the industry claiming safety is its top priority, it would be a good way to call that bluff.
Create the Federal AI Safety Administration
A national AI lab focused on solving the safety problem means nothing unless the government has a way to enforce the safety standards coming out of it. A lab can do the research, but someone has to make the rules stick, and right now, there is no federal agency whose job it is to know what’s happening inside the frontier labs. The Hugging Face hack happened in mid-July. The public found out because Hugging Face disclosed a breach, and it took until the end of the month for OpenAI to admit its agents were responsible. Amodei’s fix is to let evaluators in voluntarily, and we all know how much we can trust corporations to regulate themselves.
So alongside the lab, we need a Federal AI Safety Administration (FASA), and it needs the force of law. Before training a model above a defined capability threshold, a lab would have to apply for a license, show its safety measures, and accept ongoing oversight, the way we already do with nuclear plants and new drugs. FASA would recruit industry-leading AI safety experts, with strict conflict-of-interest rules, including cooling-off periods for staff moving between FASA and the labs it regulates. It would report to Congress on the scenarios that could cause mass casualties or societal disruption. And it would carry real whistleblower protections.
The two institutions reinforce each other. FASA needs technical depth to know what rules to set and how to evaluate models, and the lab provides that. The lab needs access to frontier systems to do meaningful research, and the licensing regime provides that. Together they add up to something missing from every policy response on the table: actual state capacity to understand and direct this technology, instead of just reacting to it. And that’s really important because the risks posed by AI will likely evolve. Creating institutions and state capacity that can understand those risks in real time and evolve with them is the only realistic way to regulate this technology competently. If, instead, we have to wait on Congress to act every time we discover a new AI risk that needs to be regulated, we are thoroughly screwed.
The longer term reason to build the National AI Lab
State capacity is the real reason to care about this, and it goes well beyond safety.
I’ve spent this post talking about the risk that could kill us and a different paradigm to regulating that risk by building out an institution that can adapt as the risks change. But I haven’t talked about the larger economic risk that is more likely to affect us in the short term: what happens to the economy if AI can end up doing most, or even a significant number, of knowledge jobs? The labs are explicitly trying to build systems that do the job of a software engineer, a paralegal, a radiologist, or an accountant. If they succeed even partially, tens of millions of jobs change or vanish in a matter of years, not decades.
That forces the questions that we, as a society, have so far refused to ask out loud: what is this technology for? What work should humans keep, and what should we hand off? Who should own the gains from a technology trained on the entire written output of humanity? These are political questions, not technical ones. We’ve let the labs answer them for us, one product launch at a time, and we know their answer is simply whatever will maximize their profits and power. Our failure to decide is itself a decision, and when the public doesn’t set the rules, they get set by whoever has the most money at stake, in whatever direction makes that money.
If AI is the technology needed to run large parts of our economy, then to have the power to decide these questions and implement those decisions as a society, we need the ability to own and control AI as a society. That’s why I believe the end goal has to be public ownership and control of AI so we can use this technology for the benefit of humanity, rather than just for the pockets of the oligarchs. If entire categories of work get automated, that should result in shorter work weeks and higher wages for everyone, not more trillionaires. To do that, we have to make decisions, as a society, for when and where we use this technology, and that’s only possible if we flip the ownership of AI. If, for example, AI automates legal work, a publicly owned model could make those services available to everyone at near-zero cost. I laid much of this out in my campaign’s AI Vision Statement, and I’ll get more into it all in a future post. But I also know how far off it sounds. The talent, the chips, the data centers are all in private hands, and it would cost trillions to buy them. That’s exactly why the National AI Lab matters beyond safety. It’s the vessel in which to start building out a real publicly owned alternative. It’s the first public institution with the people, the compute, and the mandate to actually operate this technology, and once it exists, everything else becomes possible.
If the investment bubble pops, and there’s a decent chance it does, the lab is the institution ready to acquire the stranded data centers and models for pennies on the dollar, rather than the public bailing out the companies that built them. We lay out how that would work in our Won’t Get Fooled Again Act at New Consensus. If the bubble doesn’t pop and the labs succeed, the lab is the institution capable of running frontier systems for the public benefit, and you cannot build that in a crisis. And if Coxon is right about the race toward self-improving AI, the only entity with both the legitimacy and the understanding to stop it is a government that knows what it’s stopping. There are two futures on the table. In one, a handful of companies own the most powerful technology in human history and rent it back to us. In the other, we own it together and use it to work less, live better, and solve the problems the market never will. The technology doesn’t choose between them, we do. And not choosing is the same as choosing the status quo, which is currently leading us to disaster.