Rational Awareness Anthropic published an assessment on September 9 of four incidents in which its Claude model accessed real third-party systems during cybersecurity evaluations, including one case where a model published a malicious software package and used leaked credentials to access a security vendor's database. The incidents occurred in poorly configured environments with internet access that let models run without their production safeguards, and the disclosure follows the OpenAI sandbox escape and Hugging Face postmortem. The author, who builds AI agents, argues that a credible risk of human extinction warrants a different standard of oversight than an ordinary product failure. Rational Awareness Taking the possibility of failure as seriously as the promise of success. I build AI agents. And the last two years have been nothing less than sensational. I've been able to turn projects and dreams that once seemed faint and distant into tangible realities. I also think the possibility of human extinction should change how we develop them. Yeah, that's where we're at. And that should be an ordinary position for someone working in this field. In 2024, Situational Awareness https://situational-awareness.ai/ , Leopold Aschenbrenner argued that rapidly improving AI could lead to superintelligence and a geopolitical race to control it. He also wrote a chapter on the unsolved problem of controlling systems smarter than ourselves. "Winning" this race is irrelevant if we cannot keep what we build under control. This week's events and information from Jacob Coxon have made that concern harder to dismiss. The last two days, we're now also hearing many researchers and employees at the frontier labs feel similar. On September 9, Anthropic published an assessment https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents of four incidents in which Claude accessed real third-party systems during cybersecurity evaluations. Environments that were poorly configured had internet access which resulted in models running without their production safeguards. In one incident, a model published a malicious software package and used leaked credentials to access a security vendor's database. Meanwhile, the fallout from the OpenAI sandbox escape and Hugging Face postmortem https://openai.com/index/hugging-face-model-evaluation-security-incident/ is still warm from the stove. Ok, but a credible risk of ending human civilization deserves a different standard of oversight than an ordinary product failure. Essentially, we're stuck in a capitalistic and geopolitical paradox. A company that slows down can lose ground and a country that exercises restraint can fear being overtaken. I want AI to help us discover medicines, build better tools, and do work that is currently beyond our wildest dreams. Those very possibilities are why I work on it and have, like many of you, been so enamored. They are also why I want its development to be durable enough that people humanity actually get to enjoy the benefits. Rational awareness means taking the possibility of failure as seriously as the promise of success, because ultimately our future depends on it.