cd /news/ai-safety/rational-awareness · home topics ai-safety article
[ARTICLE · art-126999] src=mouse.dev ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Rational Awareness

Anthropic published an assessment on September 9 of four incidents in which its Claude model accessed real third-party systems during cybersecurity evaluations, including one case where a model published a malicious software package and used leaked credentials to access a security vendor's database. The incidents occurred in poorly configured environments with internet access that let models run without their production safeguards, and the disclosure follows the OpenAI sandbox escape and Hugging Face postmortem. The author, who builds AI agents, argues that a credible risk of human extinction warrants a different standard of oversight than an ordinary product failure.

read2 min views3 publishedSep 11, 2026
Rational Awareness
Image: source

Taking the possibility of failure as seriously as the promise of success.

I build AI agents.

And the last two years have been nothing less than sensational. I've been able to turn projects and dreams that once seemed faint and distant into tangible realities.

I also think the possibility of human extinction should change how we develop them.

Yeah, that's where we're at.

And that should be an ordinary position for someone working in this field.

In 2024, Situational Awareness, Leopold Aschenbrenner argued that rapidly improving AI could lead to superintelligence and a geopolitical race to control it. He also wrote a chapter on the unsolved problem of controlling systems smarter than ourselves. "Winning" this race is irrelevant if we cannot keep what we build under control.

This week's events and information from Jacob Coxon have made that concern harder to dismiss. The last two days, we're now also hearing many researchers and employees at the frontier labs feel similar.

On September 9, Anthropic published an assessment of four incidents in which Claude accessed real third-party systems during cybersecurity evaluations.

Environments that were poorly configured had internet access which resulted in models running without their production safeguards. In one incident, a model published a malicious software package and used leaked credentials to access a security vendor's database.

Meanwhile, the fallout from the OpenAI sandbox escape and Hugging Face postmortem is still warm from the stove. Ok, but a credible risk of ending human civilization deserves a different standard of oversight than an ordinary product failure.

Essentially, we're stuck in a capitalistic and geopolitical paradox. A company that slows down can lose ground and a country that exercises restraint can fear being overtaken.

I want AI to help us discover medicines, build better tools, and do work that is currently beyond our wildest dreams. Those very possibilities are why I work on it and have, like many of you, been so enamored. They are also why I want its development to be durable enough that people (humanity) actually get to enjoy the benefits.

Rational awareness means taking the possibility of failure as seriously as the promise of success, because ultimately our future depends on it.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rational-awareness] indexed:0 read:2min 2026-09-11 ·