# OpenAI Teases Astra, Limits Cybersecurity Tools After Hugging Face Hack

> Source: <https://uk.pcmag.com/ai/167027/openai-teases-astra-limits-cybersecurity-tools-after-hugging-face-hack>
> Published: 2026-09-02 09:42:38+00:00

OpenAI is ready to talk more about its upcoming Astra model, although it has yet to share a launch date for its next advanced release. In a [blog post](https://openai.com/index/path-to-astra/), the ChatGPT-maker said it plans to “make Astra available soon.”

According to OpenAI, Astra will be particularly effective at identifying flaws in security systems without human assistance, confirming it believes Astra meets its own “Critical cybersecurity capability threshold."

OpenAI's definition of that benchmark says, “With the right tools and access, it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step.”

These advancements, paired with the fallout from the July [Hugging Face attack](/ai/166262/openai-oops-our-models-went-rogue-hacked-hugging-face), mean that OpenAI is taking security more seriously than ever, adding further safeguards before releasing a new model. It also plans to limit its most advanced cybersecurity features to select partners, though it hasn't revealed how it will do so.

The Astra model wasn’t involved in the Hugging Face incident, but the fallout affected all of the brand's systems.

In July, an advanced OpenAI model, working in tandem with GPT-5.6 Sol, broke containment of its sandbox by using a previously unknown zero-day vulnerability to access the internet. The models then targeted the AI platform Hugging Face, which hosts over 2 million public AI models and datasets, to help it obtain additional information.

OpenAI says, “We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity.”

Earlier this year, rival Anthropic delayed the launch of its [Mythos model](/ai/164298/anthropic-our-new-model-is-so-powerful-only-a-few-partners-can-try-it-out) due to concerns about worldwide cybersecurity. The news was largely positive for Anthropic, showing how confident it was in its major AI advancements, and it's likely OpenAI feels similar about Astra's advances.

OpenAI says it will allow a group of testers to trial its Astra tools first, but there's no word on a schedule or who it will trust those testers to be. OpenAI says, "We will continue to test these systems, share what we learn, and be clear about what remains uncertain. The models that follow Astra will demand more of us. We will take the time and do the work needed to meet that responsibility."

Not everyone is convinced by OpenAI’s new precautions, though. Former employee Yona Shavit, who now works at the OpenAI Foundation, [posited](https://x.com/yonashav/status/2094887195249168609?s=20) that if the models knew researchers were watching, it may be capable of lying through "explicit or implicit metagaming-reasoning."

*Disclosure: Ziff Davis, PCMag's parent company, filed a lawsuit against OpenAI in April 2025, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.*
