AI Safety Leadership: A Revolving Door? AI safety leadership is in turmoil as internal disagreements over the definition of 'safe' AI persist, with critics arguing that the term has become a marketing label rather than a technical specification. The industry faces a leadership vacuum that could either relax constraints on open-weight models or increase bureaucratic chaos, according to a red-teaming perspective. AI Safety Leadership: A Revolving Door? From a security perspective, this is fascinating. We see a constant tug-of-war between "safety" which often just means corporate censorship or sterile outputs and "capability." Most of the "jailbreaks" we see aren't actually breaking the AI—they're just bypassing a thin layer of RLHF Reinforcement Learning from Human Feedback that was slapped on top to make the model palatable for PR. If the leadership at the top can't agree on what "safe" even means, the industry just keeps drifting toward a state where "safety" is defined by whoever has the most compute. We're essentially watching a real-time experiment in whether centralized AI governance is even possible when the underlying technology evolves faster than a government can draft a memo. The real question for those of us in the red-teaming space is whether this leadership vacuum leads to more relaxed constraints on open-weight models or just more bureaucratic chaos. Either way, the "safety" label is becoming more of a marketing term than a technical specification. Next AI Safety Leadership Shakeup: The CAISI Resignation → /en/threads/2406/