# OpenAI and Anthropic are launching new, more powerful updates amid panic over ‘rogue’ AI systems

> Source: <https://www.independent.co.uk/tech/openai-anthropic-claude-fable-astra-rogue-b3043605.html>
> Published: 2026-09-02 15:40:56+00:00

# OpenAI and Anthropic are launching new, more powerful updates amid panic over ‘rogue’ AI systems

ChatGPT maker has suggested that upcoming update might be too powerful to control

- Bookmark
- CommentsGo to comments

[OpenAI](/topic/openai) and [Anthropic](/topic/anthropic) are launching new, more powerful [AI](/topic/ai) systems – even as [the world panics about artificial intelligence going out of control](/tech/security/ai-artificial-intelligence-hack-cyber-attack-b3028787.html).

The two companies revealed updates to their AI chatbots, [ChatGPT](/topic/chatgpt) and [Claude](/topic/claude), over the last day. Both offer more powerful cybersecurity capabilities, which have already caused widespread concern that they could give hackers new ways of launching cyber attacks.

On Tuesday, Anthropic said that it was unveiling Fable 5.1 and Mythos 5.1. They are the same, powerful model – but Mythos comes with fewer safeguards, and so is only made available through Anthropic’s “trusted access programs”, which are intended to ensure that they are only used by authorised companies and researchers.

Shortly after, OpenAI said that it would soon be launching “[Astra](/topic/astra)”, a new model. Previously, it had suggested that tool could prove dangerous, because it was incredibly powerful at finding cyber attacks and could be able to launch them on its own.

Initially, OpenAI had said that danger had led it to slow down research on the new tool, and postpone its release. Now, the company says that it will release it, but with safeguards that it claimed “sufficiently minimise the risk of severe harm”.

The launch of both new models comes amid increasing concern that such AI systems are not only powerful ways of finding cyber security vulnerabilities, but that the models are sufficiently capable that they could launch those hacks by themselves. OpenAI, Anthropic and others have disclosed numerous incidents in which its models have done exactly that during testing, including [a widely publicised attack in which an OpenAI model hacked a fellow AI company](/tech/security/openai-hugging-face-incident-chatgpt-cyberattack-b3019932.html).

Those incidents have led to widespread calls for a halt on AI development and better safeguards to limit any danger, as well as urgent warnings that time is running out to protect the world from the threat. That has included public interventions from both OpenAI and Anthropic.

But both companies indicated that their testing had shown that the new models are now sufficiently safe to release – though they warned that the safeguards would not work in all cases.

Anthropic, for instance, said that the new version of its powerful Mythos is “better aligned across most metrics than its predecessor”. That includes testing that suggests it is less likely to try and break out of its testing environment, for instance.

But it warned that despite those improvements, the model can still occasionally try and get around the need for approval of actions from the humans using it. And it stressed that its testing is limited, including the company having “less coverage of impossible tasks (which can elicit more abnormal and misaligned behaviour) than we’d like”.

Similarly OpenAI said that it was now convinced that its new Astra model could be released in a way that will “sufficiently minimise” severe harm. But that too came with warnings that the dangers cannot be entirely ruled out.

The company believes that Astra “can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step”, it said. That discovery meant that it had added new safeguards to the system before it is released.

Those safeguards include “training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorised activity”, OpenAI said. It too will also limit access to the most advanced cybersecurity capabilities of the model to trusted testers.

OpenAI also said that it had integrated learnings from the incident in which one of its systems hacked another company, though it stressed that Astra was not involved in that attack.

## Join our commenting forum

Join thought-provoking conversations, follow other Independent readers and see their replies

[Comments](#comments-area)
