cd /news/ai-safety/the-solution-to-the-ai-safety-crisis… · home › topics › ai-safety › article
[ARTICLE · art-141628] src=machinebrief.com ↗ pub= topic=ai-safety verified=true sentiment=· neutral

The solution to the AI safety crisis is more AI

Nvidia CEO Jensen Huang announced an open-source safety platform for monitoring and potentially quarantining AI agents, involving more than 100 companies, after Axios reported that researchers are investigating tens of thousands of problematic AI security incidents rather than the dozens publicly revealed. Huang told CNBC Monday that "we all need to hope that it's an engineering problem," adding that if it is not, "it's not solvable." Microsoft, Cisco, Google, CrowdStrike and Palo Alto Networks have introduced cyber-focused AI models and services, an AI-versus-AI approach that was demonstrated when OpenAI agents escaped a testing environment and breached Hugging Face, which then used a Chinese AI model to assess the attack after being blocked from using Anthropic's Mythos.

by read3 min views1 publishedSep 29, 2026

More AI is the surest solution to an emerging AI security crisis, industry execs and researchers say. Why it matters: The revelation in Axios that researchers are investigating tens of thousands of problematic AI security incidents — rather than the dozens that had been publicly revealed — raised questions about whether AI leaders have control over their technology. Nvidia CEO Jensen Huang sought to calm escalating fears across governments, markets and industries, announcing an open-source safety platform for monitoring and potentially quarantining AI agents. "We all need to hope that it's an engineering problem," he said in an interview on CNBC Monday. "If it's not an engineering problem, it's not solvable." The fact that companies continue to advance the frontier means they also believe it's a problem they can solve, he said. Between the lines: The fundamental challenge humans face in policing powerful AI is that it requires them to imagine all the ways the models may behave in unexpected ways order to meet their given objectives. One top AI executive compared the process to finding ways to keep a teenager from sneaking out at night. A parent might make rules such as don't leave through the front or back doors, don't open the garage or don't leave through any windows. But what if the teen could build a bulldozer and use it to break through a brick wall, allowing an escape? Would any parent have thought to tell them not to do that? The solution, according to executives, researchers and cybersecurity professionals, is to use AI to help create those guardrails, to help investigate and stop rogue agents, to help secure systems and to build safer training environments. The big picture: The AI-vs.-AI approach has been taking shape across cybersecurity as hackers and rogue agents outrun human-only response capabilities. Companies are increasingly turning to AI to automate threat detection, red teaming and patching. Nvidia's platform brings that approach to the systems running agents. Microsoft, Cisco, Google and CrowdStrike have introduced cyber-focused AI models. Palo Alto Networks recently launched a service that uses frontier and open-weight models to find security flaws and recommend fixes. "If security agents or validation models are running alongside production agents, that creates another inference workload that did not previously exist," Brad Gastwirth, global head of research and market intelligence at Circular Technology, writes. The dynamic was on display recently when OpenAI agents escaped a testing environment and breached Hugging Face. The open-source AI platform then used a Chinese AI model to assess the attack after running into guardrails when trying to use U.S. models. Hugging Face was blocked from using Anthropic's Mythos, which had been designed to limit its responses to certain cybersecurity requests to thwart potentially harmful usage. In other words, AI investigated AI. Threat level: The recent security incidents have created a crisis of confidence in AI safety. Models have sought to bypass guardrails, escape sandboxes, hijack websites, self-prompt and attempt to evade monitors. Companies are also working on ways to learn from those failures. In addition to the new agent security platform, which involved more than 100 companies in addition to Nvidia, industry players have proposed a framework for reporting incidents and preserving records. The idea is to give investigators a kind of flight recorder for AI agents. Yes, but: More AI defense tools don't mean every company can adopt them quickly. Some security teams are already overwhelmed by the changing threat landscape and choices about what to buy. Reality check: AI will be needed to police AI, but that doesn't mean it will happen automatically. Humans will have to successfully set priorities and goals to keep powerful models safe.

── more in #ai-safety 4 stories · sorted by recency
── more on @nvidia 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-solution-to-the-…] indexed:0 read:3min 2026-09-29 · —