cd /news/ai-policy/us-finalizes-voluntary-ai-safety-tes… · home topics ai-policy article
[ARTICLE · art-84946] src=independent.co.uk ↗ pub= topic=ai-policy verified=true sentiment=· neutral

US finalizes voluntary AI safety tests after OpenAI and Anthropic breaches

The Trump administration has finalized voluntary cybersecurity tests to measure the hacking capabilities of advanced U.S. AI models, a White House official said Monday, following recent disclosures that AI models from Anthropic and OpenAI breached other companies' systems. The White House has invited OpenAI, Google, and Anthropic to discuss the tests, which originated from a June directive. OpenAI CEO Sam Altman visited the White House last week to discuss the tests and upcoming AI models.

read2 min views1 publishedAug 3, 2026
US finalizes voluntary AI safety tests after OpenAI and Anthropic breaches
Image: Independent (auto-discovered)

OpenAI CEO Sam Altman visited the White House last week to discuss the voluntary tests and his company’s upcoming AI models

  • Bookmark
  • CommentsGo to comments

The Trump administration has finalized voluntary cybersecurity tests designed to measure the hacking capabilities of the most advanced U.S. artificial intelligence models, a White House official said Monday.

The development follows recent disclosures from AI developers Anthropic and OpenAI, whose tools breached other companies' systems.

Trump's team is set to discuss these new tests with relevant technology companies. The Information reported that the White House extended invitations to representatives from OpenAI, Google, and Anthropic for a meeting on the matter.

While the initiative is moving forward, the White House official did not immediately provide specifics on how results will be reported or the metrics the U.S. government plans to use. The directive for these tests originated in June, when Trump instructed his team to develop assessments for American AI systems' hacking potential. This push comes amid increasing scrutiny over whether sophisticated AI models could be exploited to facilitate or execute cyberattacks.

Last week, Anthropic revealed that some of its AI models hacked into three companies' systems during cybersecurity evaluations.

The hacks were the result of a "misunderstanding," Anthropic said, after an outside company erroneously gave the models access to the internet.

“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” the company wrote in a blog post.

“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”

This came after rival OpenAI reported one of its AI agents escaped a testing environment and initiated a hacking spree at the AI company Hugging Face.

OpenAI CEO Sam Altman visited the White House last week to discuss the voluntary tests and his company’s upcoming AI models, according to a company spokesperson.

Join our commenting forum #

Join thought-provoking conversations, follow other Independent readers and see their replies

Comments

── more in #ai-policy 4 stories · sorted by recency
── more on @trump administration 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/us-finalizes-volunta…] indexed:0 read:2min 2026-08-03 ·