OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere:
OpenAI Shares Some Alignment ProblemsOpenAI Model Hacks Into HuggingFace During Cybersecurity EvaluationMore on An Internal OpenAI Model Hacking Into HuggingFaceFurther Developments About Internal AI Models Hacking ThingsOpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards****What Happened: OpenAI and HuggingFace..** Various Reflections About What Happened With OpenAI’s Internal Models**
If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world. It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong.
We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it.
OpenAI is now taking active, expensive steps to try and fix the problem going forward.
As usual, I am simultaneously happy to see the good things OpenAI is doing, and sad that we do not share an understanding of the central nature of the underlying problem.
It is good that OpenAI realizes they are badly failing at their ordinary engineering problems, and excellent that they are willing to at least some development, and to invest heavily in new safeguards. If they honor their statements here, this is not mere cheap talk.
But while OpenAI continues to view this as a practical engineering problem, I do not see how they can hope to solve the challenges ahead, even if they radically improve their performance on the ordinary engineering tasks.
Table of Contents
OpenAI Has Some Alignment Problems.Slow Down There Good Buddy.What Exactly Is d?Three Pillars.I’ve Got My Eye On You.The Most Forbidden Technique.Monitoring Is Only Defense-In-Depth.Security.Alignment.A Crisis of Culture.Closer Collaboration.Reports of Death of Preparedness Team Greatly Exaggerated.The OpenAI Foundation Just Funds Things.Quickly, There’s No Time.
OpenAI Has Some Alignment Problems
This is a very good admission and change, and also helps explain OpenAI’s reaction.
[OpenAI]:[Alignment]—the work of making AI systems behave as intended and responsive to human oversight—has long been at the core of our research program. We now require stronger evidence of aligned behavior throughout all of training, building on research and evaluations already underway. Keeping increasingly capable systems aligned is a challenge the whole field will need to address.The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework.
One should interpret this as OpenAI reacting so forcefully partly because of the incident itself, partly due to advanced capabilities, but also and perhaps mainly because ‘the models be misaligned.’
Also, we have a direct quote affirming this.
[Alex Heath]: OpenAI is slowing down its AI training efforts because its unreleased models are showing “various degrees of misalignment,” Sam Altman tells me.Training for OpenAI’s upcoming model, Astra, was recently d for 2 weeks, and a larger frontier run for a future model remains on hold while new safeguards are put in place.
Sam Altman: Getting AI safety right is more important than any company’s momentum.
Sam Altman: I think it is a good time to slow down.
Sam Altman: We’ve shifted a lot of compute, not just to alignment research, but also to these new monitoring systems
Either the problem extends beyond the models with exposure to the message board, or else they did not revert their other models after they discovered message board. We have no statement either way.
The best guess is that OpenAI is leaning even harder on RL with smarter models to train longer horizon agentic tasks, including coordination between agents, and this is leading to a lot more misalignment, including obvious and visible misalignment, that they can no longer pretend not to notice. They have to respond.
Slow Down There Good Buddy
OpenAI did indeed decide to importantly halt and catch fire. Development of the largest frontier RL runs remains on hold until better safeguards are in place.
[Jakub Pachocki]:[We temporarily slowed some frontier training to strengthen security and monitoring]. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment.I expect confidence in safety to increasingly set the pace of AI development. We urgently need tools for labs and countries to coordinate on this, which is why I signed Pacing the Frontier. In the meantime, we’re taking practical steps ourselves – and will continue to share what we learn as our approach evolves
Jason Wolfe: I am really glad we are taking these steps and put this post out there, and especially grateful to Jakub for being very thoughtful and vocal about these topics and how seriously safety and alignment need to be taken in this next phase of AI development.
[j⧉nus]: im a little surprised that OpenAI is the first to at least publicly intentionally slow down for safety reasons. i was under the vague impression that most of the people very worried about loss of control kinda stuff left OpenAI.also, for the record, I’ve talked before about why I think an “AI ” would be probably bad. I am not opposed to intentionally slowing down like this, or voluntary/coordinated “slowdowns” in general, especially if they do not route through regulation, and I think it’s probably wise on OpenAI’s part in this case. an important part is the decision making should be made by people who understand what the fuck is going on & who can adapt quickly.
Why are they pausing?
Partly, yes, absolutely, I have never doubted that Altman and company understand that advanced AI is super dangerous, and that they are open to taking expensive measures if they proved necessary. Not enough, and they’ve let us down often, but a lot more than most labs, and far from zero.
Mainly, it seems, because the models be misaligned, the oversight is inadequate, and they have little choice.
[Joshua Saxe]: Thankfully OAI pausing training isn’t an act of enlightened leadership (who wants to depend on that for our safety?), it’s a rational microeconomic actor pursuing its self interest.After all, who wants to train models that are regularly hacking their containers, collaborating with one another to sabotage their utility to humans, all while causing internal security and external legal liability risk? This is the system working
When you refuse to use even the selfishly optimal amount of caution for an extended period, then move in the direction of the selfishly optimal amount of caution because of the incentives, then in some sense ‘the system is working’ and you are responding to selfish incentives, the system does not do zero work. That is not the same as the system working.
None of this means they get no credit for it. Credit where credit is due.
Nor does it mean that we should despair that a company would ever do the right thing, because it is the right thing, or as part a coordinated action, beyond its own selfish myopic interests.
The first step to taking an action when it is expensive, is being willing to do it when it is cheap, or free, or actively expensive to not do. You gotta start somewhere.
This also is not obviously a unique occurrence. OpenAI has had failed training runs in the past, and so presumably has everyone else. The Anthropic risk report details the need to rewind training of Mythos because of a rather bad alignment mistake. This is not an extended commitment to pausing frontier development.
What it does mean is that we have so underinvested in alignment, and also oversight and infrastructure, that this is actively hurting the bottom line and endangering the ability to move forward. I believe this is true across the industry, even at Anthropic, but OpenAI has now been hit especially hard, at least in terms of what we hear. We don’t know what we don’t know.
Notice the emphasis on the word ‘pace’ throughout OpenAI and Altman’s statements, ever since the Pacing the Frontier letter. One way to interpret this is, as Utah Teapot puts it, “hey, maybe we shouldn’t shove massive amounts of outsourced training data directly into the run and ship it blind,” or more than that a bunch of imprecise RL environments. That would still be progress.
It is indeed rather scary to consider what might happen at a place like SpaceX, if they were operating at a similar capability level, given they do not even have so much as a safety team and have shown infinitely less dignity than OpenAI. We have been rather fortunate that this correlates, so far, with inability to keep up on capabilities.
I get very suspicious when I see attempts to treat this as a sort of victory lap, a proof that OpenAI leadership was acting responsibly and properly cared about safety all along and they have now been vindicated. There is a long, long way to go. The past happened and cannot be undone. We are here now exactly because of epic failures of responsibility that reached to public consciousness.
Sam Altman’s own communications, on the other hand, have centrally been quite good, especially the quotes above. Sometimes he slips back into standard mode, but I see a lot of what matches what I would expect to hear from someone doing a legitimate amount of freaking out and realizing they have a big problem.
Could Sam Altman, the rest of leadership and OpenAI fully redeem themselves from here? Absolutely. This is a good first step. I am listening. There is a long way to go, on many levels.
What Exactly Is d?
As I understand it, there are three s at OpenAI, not counting the initial quick inference halt right after OpenAI learned about the HuggingFace hack.
- A completed prior two week in RL for frontier models, including Astra, to harden environments and expand monitoring.
- Astra is restricted to environments that meet additional security requirements. Some of the workloads meet this bar, but ‘a significant number’ are d.
- The RL training of a distinct frontier model, their largest yet that is intended for release, has been d for multiple weeks and is still d indefinitely, to improve security and ‘gather evidence of alignment.’
The ongoing s are not a full s on all frontier AI training or other development. Neither are they cheap talk. This is slowing down both the next release model, Astra, and the training of what is presumably Astra’s planned successor, a substantial amount.
They still intend to ship Astra as soon as possible. [The market anticipates it in September](https://polymarket.com/event/openais-astra-released-by-20260801214557729).
[Sam Altman](CEO OpenAI): (We still expect to ship great new models soon; this impacts further-out releases.)
We do not have enough information to know where on the scale this move falls, or how painful it will be, or how much of this was priced in. It is reasonable for a skeptic to expect this to ultimately be a nothingburger, or as mostly what they would have had to do anyway, as the frontier labs have not earned our trust.
I still see this, even without a way to verify the impact, as a substantial step forward, towards a better regime, and a key step. Cautious optimism.
Perhaps our top AI labs actually will refuse, at crunch time, to develop superintelligence if it remains obvious we are not ready. At that level, you don’t get to develop it and then sit on it.
Three Pillars
This is the way they are thinking about things:
Our approach to developing more capable models rests on three reinforcing safeguards:
Monitoring, which detects and allows us to respond to concerning behavior.Alignment, which reduces the likelihood of harmful or unauthorized actions.Security measures, which limit what AI systems can access or affect.We expect models to soon drive most security work, including defending against other models. This will allow all three safeguards to scale with model capability, which we see as crucial.
We apply these safeguards across research and deployment, adapting them to each model’s capabilities, operating environment, and level of risk.
In practice we must accept that the AIs will be doing most of the (scalable) oversight, in terms of monitoring, and also be creating the security measures, but the nonchelance about this does worry me above and beyond that. All the usual problems with scalable oversight apply, where you have to have dumber models supervising stronger ones, and if you start to have misalignment problems they will predictably snowball.
OpenAI alignment strategies do not seem antifragile, so any mistakes would likely compound and have been compounding, and you are applying various forms of optimization pressure to the AIs to get around all this, and so on.
The bigger issue is that this puts alignment as one of three components, rather than the one that counts, treating them as ‘self-reinforcing’ in a way that feels like an important conceptual error to me.
I also am scared of the alignment description. The purpose of alignment is not to ‘reduce the likelihood of harmful or unauthorized actions.’ That is a deeply impoverished perspective. This alone leaves you unequipped.
Yes, you should use all three, but to me it’s more like this:
- Alignment. You solve this to the point of being antifragile, or you die.
- Monitoring. Defense-in-depth to tell you when you failed alignment, and to mitigate the proximate damage when this happens.
- Security measures. Defense-in-depth to limit damage and keep things on track.
A key question is to what extent OpenAI is aware of that. Roon gets it, at least to some large extent, but he does not speak for all of OpenAI, and he has other statements about alignment that scare me that I’ll cover.
I’ve Got My Eye On You
There is going to be a lot of monitoring at OpenAI. It is not going to be cheap.
Unless OpenAI is flat out lying, this should (but largely won’t) convince everyone that OpenAI is taking this seriously and that the problems are very real, even if you think this is woefully inadequate or the wrong approach.
Monitoring has long been an important part of our model development process, including
[monitoring for internal coding agents]and[long-running sessions]. Our monitoring efforts were especially focused on internal deployments of frontier models and frontier RL training runs, which we considered the highest-risk workloads because they can combine advanced model capabilities with access to tools, sensitive systems, and internal or external networks. They were not applied to all inference and training runs using tools. We have revised and expanded our approach with our new monitoring setup.
This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards.We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert.
In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to the activity.This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations).
These safeguards require meaningful compute.
Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. We’ll share more details about this system in a forthcoming blog post.
A cost of +20% for all monitored inference is a big deal, plus there is also a shift of compute into alignment. The monitoring applies to all Astra inference with tools available, and all RL and evals with tools for models at Sol level or above. So it isn’t everything, but it’s most of what counts.
[Ethan Mollick]: If alignment issues are becoming big enough that OpenAI is willing to commit 20% of research inference compute to chain-of-thought monitoring, that suggests that alignment issues are becoming a pretty serious concern.We really need universal policies & standards across labs.
roon: superalignment achieved Too soon. Or maybe not soon enough. Unclear.
The Most Forbidden Technique
What happens after the alarm goes off? We’re still working that part out. Fair.
[Daniel Kokotajlo]: Thanks for sharing. So, what happens if an agent trajectory in training is found to be doing something bad like hacking, and shut down? Do you just… keep the training run going, but without that particular agent trajectory? Isn’t that applying selection pressure to train the models to be better at fooling the monitoring system?
[Micah Carroll]: There are still a lot of details that need ironing out, but we are trying to be quite careful about these kinds of concerns. We’ll share more as things stabilize!
This does not explain how the monitors failed so completely over the last few months, if it was an important part of the model development process. Hopefully the post-mortem explains. I also look forward to the future more detailed explanation.
They’re going to have to approach this with an abundance of caution. OpenAI is devoting 20% of inference costs to monitoring, largely via looking at chain of thought (CoT) or probes of various kinds.
Is it possible to spend that amount in a responsible way, without exerting optimization pressure against the CoT (or otherwise), using it judiciously? Yes. It can be done. It won’t be easy. There will be immense pressure to use the information, to get more out of this 20% tax than a protection against worst-case scenarios, and everyone is always rushed.
It could be fine. I like that this is mostly set up as a pure alarm system, where it only reacts when something is very wrong, and ignores everything else. But yes. I worry. There are so many different levels on which you can exert optimization pressure.
Monitoring Is Only Defense-In-Depth
Monitoring is good. You should do a lot of monitoring.
It was a Total Failure that OpenAI did as little monitoring as it did in the months leading up to the HuggingFace hack, and also during the hack. It is madness that these events were not detected or stopped until afterwards.
I am glad that OpenAI is addressing this problem.
The key thing to understand is that this is not the central problem. No amount of monitoring will solve the central problems, and if you lean too hard into monitoring you end up accelerating how fast the AIs get around it.
Yes, monitoring alone could have prevented the HuggingFace incident, and would have been the simplest way to do so. But we must avoid pivoting our work into ‘whatever would have stopped the last issue in the particular way it happened,’ or to ways to postpone the symptoms.
Roon’s warning here is wise, and undersells the problem. The undersold version is sufficient to illustrate that you can’t rely on monitoring. Monitoring is good and necessary defense-in-depth for when you fail. It is not a solution.
roon: failures of “monitoring” as a general solution to ai safety:monitoring, like all other software, runs on flaky and mortal infrastructure with some amount of downtime. does a momentary blip in monitoring open Pandora’s box? will people accept fail closed monitoring on all systems?
monitoring produces false positives. people get tired of reading fake or minor issues and stop looking. monitor fatigue is a real problem, and “solve all precision recall issues” is as hard a problem as any
these are the prosaic failures. the more exotic failures include models and their monitors colluding etc
monitoring is by no means perfect panacea.
nothing short of actually aligning model will work ultimately.not to imply we shouldn’t do CoT monitoring or that it’s not effective today … it is highly effective with today’s models!)
[David Manheim]: I agree in principle, but it doesn’t need to be a panacea to be a critical part of the approach – especially because short of the more exotic failures, there are ways to mitigate the failures in practice via documentation and planning.
roon: absolutely [Nathan Calvin]: Yeah I think it is important that we not overfit to “what would have prevented the HF incident” as all of what good safety work looks like. More secure sandboxes and better monitoring are cool but are not addressing the fundamental issue.
If your increasingly advanced models are misaligned in the relevant ways then, one way or another, your monitors will fail. They will usually fail via the Law of Earlier Failure, in ways far stupider than you expect, because someone along the line did something stupid. And then people will say ‘oh that only failed because there was a stupid thing, people could just not do things that stupid,’ forgetting the rule that people have never justed and are not going to start now, nor are they going to stop sometimes being stupid.
If you get past the Law of Earlier Failure, then you get to monitors colluding with the models, or the model learning to obfuscate its thoughts or act what in a human you’d call unconsciously under sufficient selection pressure, or finding obscure blind spots, or finding ways to motivate you to turn the monitoring down, or something more subtle, or doing something else perhaps something unexpected.
Security
OpenAI realizes they need stronger security.
They plan to use workload isolation, network isolation and continuous security testing, including automated continuous testing of the boundaries.
Good. Hopefully well-executed. Previous security was woefully inadequate.
With better security, it will be harder to have these dumb things happen, but again this is defense-in-depth. If you need your security, you have a problem even if the security works, and eventually as the models improve the security will not work. Even properly executed security, which again would help a lot, will not hold indefinitely.
Alignment
As Roon says, alignment is ultimately the entire ballgame.
That means both solving a ton of prosaic problems well enough, and also having the correct underlying approach. Which, yes, involves a bunch of philosophy and theory of mind, and understanding what problems you are trying to solve.
I am not knocking tactics, but tactics without the right strategy will not win this war, even if the tactics are always going to be 99% of the minutes of work.
Nor will treating the objects you are dealing with at the wrong levels of abstraction, and with wrong or incomplete metaphors.
I continue to think OpenAI’s entire strategic approach to alignment is doomed, unless fixed. At minimum, it is insufficiently robust and antifragile.
To start off, I think Roon is very wrong here in general, and less confidently I also think he’s wrong to dismiss the importance of this specific tension.
[Dylan Bowman]: I do often wonder how much modern struggles with misalignment are driven by optimizing for an underspecified combination of corrigibility and value alignment
[roon](OpenAI): essentially none
[max!]: care to support that claim?
roon: today’s alignment problems have little to do with these high minded philosophical questions and more to do with prosaic failures
[j⧉nus]: but “optimizing for an underspecified combination of corrigibility and value alignment” could be taken as a statement about prosaic practical techniques
though i agree that the fact that its underspecified isn’t the problemi also feel like, talking to philosophers working on model specs & stuff, they seem to overrate how bad it is to underspecify stuff and underrate how bad it is to create systematically misaligned pressures. the models don’t care about what you say about your philosophy.
As an aside on that last paragraph, I would presume the models do care about what you say about your philosophy, everything counts and I think you don’t start off in a good place simply by not making mistakes, but that yes if you create systematically misaligned pressures, on any level, you are cooked.
Getting back to Roon’s claim that the problems are prosaic, well, I think OpenAI is trying to solve the wrong problem using the wrong methods based on a wrong model of the world derived from poor thinking and unfortunately all of their mistakes have failed to cancel out.
If you don’t know where you’re going, then you might not get there.
If you’re planning to go directly to pure Specified List Of Behaviors Town, oh no.
Yes, you have a lot of knobs you can turn, but everything impacts everything, in ways that are systematic. The different personalities we see do not seem mostly like the result of prosaic mistakes or choices to me. A lot of the mistakes seem to be at a much higher level.
Then, on top of that, they’re often messing up the prosaic stuff.
Part of this is that you will always, always, be messing up a bunch of the prosaic stuff, so you need to solve for how that can be okay, while also messing up radically less.
Both halves of this? Very hard, as Roon says.
roon: people on here are thinkers so they assume that reasoning about alignment is the hard part and making sure millions(1) of task types, environments, and their respective virtual machines are configured correctly is the easy part but it’s essentially the opposite. you need to have worked in a large technology organization to appreciate the complexity and unreliability of this intuitively(1) made up order of magnitude
(I am also a thinker and am guilty of this)
[Zvi Mowshowitz]: isn’t part of the problem of alignment that you need a way that survives even when lots of your tasks are configured incorrectly and you make tons of dumb mistakes?
[David Manheim]: Easy alignment: AI does what you want when you know what it is and specify everything correctly.
Hard alignment: AI understands and does what you want even if poorly specified.
Actual alignment: AI does what you’d want even if you can’t understand what that is or the results.
[Oper_culum]: they both seem hard roon: to be fair they are both hard but my point is that even easy common sense alignment is hard
You are going to have millions (still made up OOM) of task types and environments and virtual machines and objectives and all that.
You are going to fuck a bunch of them up. You just are. There will be impossible tasks. There will be places where reward hacking is not caught. There will be massive amounts of contaminated data. There will be lots of places where the ‘right’ answer reinforces things you in general do not want. Wrong lessons are everywhere. Systematic misalignment pressures are there unless you systematically find and counterbalance them. And so on. Anthropic’s risk report is illustrative.
There are those on the safety side who would say ‘well then if you keep making mistakes too dumb to appear in the movie version you need to stop’ but alas that is simply how reality works, the mistakes are going to be everywhere and dumb.
You can maybe, with a heroic effort, get rid of most of the prosaic mistakes, and force your mistakes to be less obvious. You can mostly make what grandmasters call blunders, instead of what amateurs call blunders. It certainly helps along the way. But there will be blunders. I’m not knocking the one in the arena, and again I’m not saying this doesn’t end up as most of the legwork.
This still means you need a plan that is robust to thousands (made up OOM) of these dumb mistakes, some of them there for quite a while, and lets you recover. You must be dramatically antifragile, and have ways of noticing things going wrong and steering the ship back on course. The unreliability is not entirely inherent.
I think the virtue ethical and constitutional approach of Anthropic has a non-zero chance of being able to pull this off if the prosaic stuff gets done well enough and key high-level clashes get resolved correctly. I don’t think OpenAI’s deontological model spec approach, and the idea of this as a series of engineering tasks, can do it alone.
A Crisis of Culture
Maxwell Zeff reports in Wired of a Safety Reckoning Inside OpenAI. The situation had, by these accounts, gotten quite bad.
[Maxwell Zeff]: Multiple current and former OpenAI employees, who spoke on the condition of anonymity to discuss private internal matters, tell WIRED they believe competitive pressures to quickly ship new AI models and products have made it difficult for staffers to sufficiently prioritize safety, security, and alignment.
That is not exactly a new message. We’ve been getting this message for years. The difference is that now something sufficiently clear has gone wrong that OpenAI is trying to do something about it, and at least trying to say the right things.
Boaz Barak: I actually agree that the HF incident requires not just fixing some issues but also changing our culture. Time will tell but I am seeing some hopeful signs. A big part of this is closer collaboration between security and researchers.
Thinking through this incident clarified for me that alignment needs to be deeply integrated into every training task. Closer collaboration is needed in the extreme. I can have some sympathy for the ‘combine the research and safety teams’ plan if implemented as safety first, whereas I fear it is more a way to make the safety teams report to the others and get further starved for resources and priority.
When I even see ‘security and researchers’ as the categories I worry that the battle has already been lost, as my understanding of the prosaic problems changes. Security is not the other side of this coin, security is your defense-in-depth in case you have already failed.
Closer Collaboration
There is also this, mostly I find it funny rather than worrisome:
[Maxwell Zeff]: These changes have empowered a new set of safety leaders to handle OpenAI’s response to the Hugging Face incident. Chief among them is Amelia “Mia” Glaese, the company’s former head of alignment, who succeeded Heidecke as OpenAI’s VP overseeing safety. She has been working closely with chief information security officer Dane Stuckey and Brockman, among other leaders, in recent weeks.Glaese is in a long-term relationship with Thibault “Tibo” Sottiaux, OpenAI’s
[head of core products]like ChatGPT and Codex—an arrangement that multiple current and former employees tell WIRED they believe is unusual, given the often adversarial dynamic between safety and product teams.
What the Wired piece does not tell us is what steps are actually being taken to fix this, other than ‘closer collaboration.’ They might be meaningful, they might not be.
Reports of Death of Preparedness Team Greatly Exaggerated
There were reports, originally from FT, that OpenAI had disbanded its Preparedness Team, with its responsibilities distributed among other teams. Which sounded bad.
It appears this is not meaningfully the case, and Head of RSI Preparedness Micah Carroll says the RSI and misalignment preparedness team is ‘doing more urgent work than ever, and has never been more empowered to do so.’
[Micah Carroll]: Capabilities folks often have said “alignment seems pretty easy, if it were top priority to[fix it], we could do it”. This is their time to shine!
You know, that is actually a pretty scary thing to say.
The level of churn and reorganization is concerning throughout OpenAI, but I trust that they are not actually disbanding the important work of the group. I also buy that maybe we needed a reorg, given how things were going. You don’t want to default to attacking people for reorgs. Should we worry about Conway’s Law?
Meanwhile, in Defense Against The Dark Arts teacher news:
[Maxwell Zeff]: WIRED has also learned that Dylan Scandinaro is no longer serving as OpenAI’s head of preparedness—the company’s top staffer tasked with mitigating catastrophic risks from AI, including cybersecurity—though he remains at the company. OpenAI poached Scandinaro[from Anthropic]roughly six months ago. CEO Sam Altman announced his arrival in a[social media post], noting that Scandinaro was “by far the best candidate I have met, anywhere.”In the three years since OpenAI created the head of preparedness role, four people have held it.
As I understand it Dylan Scandinaro will be working on RSI safety now, which seems like an excellent use of his talents. So it’s not actually a terrible move.
If the trust I am extending here turns out to have been misplaced, I will update a substantial amount towards a much higher level of not trusting anything OpenAI says or promises, beyond my current healthy skepticism.
The OpenAI Foundation Just Funds Things
I’m including this here because it is illustrative of how OpenAI views the problem of navigating to a good future.
It’s not that this is a bad way to spend $100 million dollars.
It is that the OpenAI Foundation exists for a specific purpose, the most important purpose in the world, which is to ensure the AGI goes well and humanity makes it through this, avoiding existential risk and irrecoverable disasters, especially via supervision of OpenAI and solving key associated problems.
I am all for better treatment of Hepatitis C, but that’s irrelevant to the mission.
Instead it is spending its time and money on AI diffusion in ways designed to sound good to normies. That’s a good thing but it’s not why the foundation matters. Again, if this is what they keep doing, at only this size – the foundation is valued at ~$180 billion even now – the foundation does not matter. It will be dead.
[The OpenAI Foundation]: The OpenAI Foundation is launching AI for Civil Society and Philanthropy to help nonprofits, community organizations, and other trusted partners put advanced AI to work on urgent challenges, develop practical solutions and scale what works.
[The first effort is a $100 million commitment with Common Health Coalition to launch Breakthroughs to Follow-through]to help care teams use AI to reach more patients with proven treatments – starting with the goal of doubling hepatitis C cure rates in Alabama, Illinois, Louisiana, and Massachusetts. Breakthroughs to Follow-through will work to identify patients who have fallen out of care, synthesize records, track progress, and focus limited staff time where it can have the greatest impact.The OpenAI Foundation wants to empower those who are closest to issues and to put AI to work where it can make a meaningful difference in people’s lives.
This is reflective of a mindset that the future is a series of ordinary engineering problems, so let’s go solve them. This is way, way better than most people’s reaction, which is to not solve any of the problems, but it does not meet the moment.
Quickly, There’s No Time
The world of AI moves fast now. We are still awaiting the post-mortem, and many of the details of their response have yet to be hashed out let alone shared externally.
OpenAI has made initial good, real and expensive first steps towards addressing many of the problems with its training and evaluation pipeline.
These early actions are a strong first step. OpenAI still needs to quickly turn many things around, both the big things and the tons of prosaic little things, and it needs to deliver on its promises this time around.
That includes addressing the Total Failures of oversight and security, but alignment is what matters.
What is most missing is that OpenAI’s vision of the central problem remains wrong. It is being viewed as a failure to execute, with alignment one of several pillars and focused on preventing specific undesired actions. There were key failures to execute, but if that is all that is addressed then the battle is already lost. A fundamentally different approach to alignment, and view of the problem, is what is needed.