If you tell me the technology you are building could threaten humanity, my next question is going to be pretty straightforward: What are you building to prevent that? I want to see the work. Show me the dangerous behavior you are trying to prevent, the control designed to prevent it and the evidence that the control holds when somebody tries to break it. Tell me who can stop the system when the test fails.
That is a conversation I can do something with.
I remain bullish on AI. I want the discoveries, the productivity and the opportunities. But enthusiasm does not relieve the people building and deploying these systems of the obligation to explain how they intend to keep them under control. The more consequential the capability, the more substantial that explanation needs to be.
That is the starting point for my new Techstrong special report, AI Extinction Risk: What Can We Do That Actually Matters? The report examines the warnings, their limitations and a practical safety program. The question driving it is simple: What protects us when someone refuses to behave responsibly?
Read the full special report: AI Extinction Risk. You do not need to settle on a percentage chance of human extinction to recognize a serious engineering question. A system taking an unauthorized action deserves investigation. So does a safeguard that fails under unfamiliar conditions. Neither, by itself, establishes that humanity is approaching its final software update.
We need to be careful about that distinction. Describing a catastrophic scenario does not establish its probability. Dismissing every control failure as a chatbot doing something weird does not make the failure disappear either. The useful work begins with identifying what happened, what authority made it possible and what would have interrupted it.
For an enterprise deploying agents, that should sound familiar. Imagine assigning an AI agent to investigate a software problem. It needs selected logs, relevant documentation and a place to test a proposed repair. Why would that assignment automatically entitle it to production credentials, unrestricted access to other services or the ability to expand its own permissions?
Being smart enough to propose an action does not establish authority to perform it. An excellent diagnosis is still not permission to operate.
That distinction becomes especially important when we move from assistants that answer questions to systems that carry out assignments. We should ask how much autonomy the job actually needs before celebrating how much autonomy the product offers.
The report develops three connected areas of work: building safer AI, constraining dangerous use and protecting potential targets. They belong together because each addresses a problem the others cannot solve alone.
Better model behavior matters. Researchers should investigate whether protections survive unfamiliar attacks, modifications and changes in operating conditions. A successful evaluation deserves attention, but its conditions belong beside its result. Testing one configuration does not certify everything somebody might later build around it.
The surrounding controls matter just as much. Permissions, monitoring and the ability to revoke access should not depend entirely on the cooperation of the agent being supervised. I would be very uncomfortable with an employee who could silently give themselves additional privileges, erase the audit trail and disable oversight. A persuasive conversational interface does not make that arrangement more sensible.
This is also where “human at the helm” has to mean something operational. The person responsible needs enough information to understand the proposed action, enough time to consider it and actual authority to intervene. A button that says “Approve” is not evidence of meaningful oversight if the workflow makes informed refusal practically impossible.
And then there is the part that receives too little attention in conversations centered on the laboratories: the people facing an AI system they do not control.
A hospital cannot assume that an attacker will preserve a model’s safeguards. A business cannot make its security depend on a criminal following the developer’s acceptable-use rules. The target needs defenses that can withstand harmful actions regardless of which tool produced them.
That means examining access, dangerous dependencies, verification and recovery. Can a consequential instruction be checked through a separately established channel? Can compromised credentials be revoked? Can essential operations continue when one system fails? Has anyone rehearsed the recovery?
These measures do not establish that we can prevent every catastrophic scenario. They do give technology leaders concrete work to assign and results to examine. The report also distinguishes software defenses from risks involving physical systems or public health, where the relevant domain experts need to help define and test the protections.
Of course, “accelerate safety” can become another empty slogan. Put it on a conference banner and it sounds wonderful. Put it in a budget and an engineering schedule, and we begin to find out whether anybody means it.
Who owns the control? What resources do they have? Who can challenge the result? What happens when the system changes? What evidence would make the organization withhold a particular deployment?
Those questions should be part of the product conversation, with people empowered to act on the answers. An assessment that cannot affect a release decision offers very limited comfort.
The full report goes further into the research, the limits of current safeguards and the evidence organizations should request. It is intended to move the discussion from alarming possibilities toward work whose results can be examined.
I do not expect a promise that software will never fail or people will never misuse it. I expect the design to account for both.
If an AI builder wants me to take an extraordinary warning seriously, I am willing to listen. Then I want to meet the people responsible for the protection, understand what they are testing and see what happens when it fails. That is how concern becomes useful.
Read the complete Techstrong special report: AI Extinction Risk: What Can We Do That Actually Matters?