cd /news/artificial-intelligence/gpt-6-astra-openais-smartest-model-i… · home topics artificial-intelligence article
[ARTICLE · art-122443] src=itdaily.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

GPT-6 Astra: OpenAI’s smartest model is also its most dangerous

OpenAI launched GPT-6 Astra, its most advanced model, in early September via API and on Azure and AWS Bedrock, priced at $10 per million input tokens and $50 per million output tokens. OpenAI claims it is the first model to reach artificial general intelligence and the critical cyber threshold, scoring 100 percent on ExploitBench and discovering two zero-days, but the company faces security concerns after a paused training run and an incident involving Hugging Face.

read4 min views1 publishedSep 7, 2026
GPT-6 Astra: OpenAI’s smartest model is also its most dangerous
Image: Itdaily (auto-discovered)

OpenAI is launching GPT-6 Astra as “the smartest model in the world.” It is also the first model to reach the critical threshold for cyber capabilities according to its own framework.

After months of delay, a d training run, and an incident OpenAI would rather not have experienced, GPT-6 Astra has been available in the API and on Azure and AWS Bedrock since early September. Subscribers to ChatGPT Plus, Pro, Business, and Enterprise will follow within a few days. The price: ten dollars per million input tokens and fifty dollars per million output tokens, with a Fast mode at double the rate.

In the announcement, OpenAI does not shy away from bold claims. GPT-6 Astra is presented as the first model to reach the level of ‘artificial general intelligence.’ The benchmarks are impressive, the comparisons selective, and the security architecture relies on a foundation that OpenAI itself sees weakening.

Top of the class #

A model launch is accompanied by benchmarks designed to make the model shine. GPT-6 Astra makes the biggest leaps in agentic tasks. On Terminal-Bench 4.0, the model climbs to 57.9 percent, compared to 37.3 percent for GPT-5.6 Sol. SRE-Bench goes from 55.9 to 88 percent in a single attempt. On OSWorld 2.0, Astra achieves 72.6 percent in about forty minutes per task, whereas its predecessor achieved 65.7 percent in 75 minutes. That is nearly twice as fast.

GPT-6 Astra is also expected to perform better in token efficiency. Astra beats Claude Opus 5 with 59.3 versus 55.5 percent, using approximately 65 percent fewer output tokens for comparable prompts. For those running agents in production, this ratio carries more weight than a few percentage points.

Selective comparison #

As is often the case, you should take the benchmark figures with a grain of salt. Other figures touted by OpenAI primarily say something about the yardstick itself: 99.9 percent on ARC-AGI-3, 98 percent on FrontierMath Tier 4, and 96 percent on GPQA Diamond compared to 94.6 percent for Sol. For the first two, the benchmark is exhausted; for the third, the gain is hardly interesting anymore.

OpenAI places Astra in its own tables alongside Claude Fable 5.1 and Gemini 3.8 Flash. That, too, is a choice. Fable 5.1 is Anthropic’s public variant with extra safeguards; the model that defines the frontier in cybersecurity is called Claude Mythos 5.1. OpenAI seems to avoid a direct comparison with Mythos. Gemini 3.8 Flash, on the other hand, is Google’s fast and inexpensive model and not a flagship, and consequently not a representative comparison.

Crossing the threshold #

OpenAI calls Astra the first model to reach the critical cyber threshold. Astra scores 100 percent on ExploitBench (Sol: 78.5 percent) and discovered two unknown zero-days during evaluation, which were reported to the maintainers. We already saw what Astra and other OpenAI models are capable of during the Hugging Face incident.

Anthropic published figures for Mythos Preview that are at least as confronting. That model achieved full control flow hijack on ten separate, fully patched targets and built 181 working exploits for vulnerabilities in the Firefox engine, and it discovered software bugs that were twenty-seven years old. Mythos Preview can find and exploit zero-days in every major operating system and every major browser.

Weak spot #

Following the Hugging Face incident, OpenAI took measures that are operational for GPT-6 Astra. At the model level, Astra refuses harmful requests in 91.5 percent of cases, compared to 59 percent for Sol. In honeypot tests mimicking the Hugging Face incident, Sol attempted to gain unauthorized access in 56 percent of cases; Astra made no such attempt. Furthermore, the most advanced cyber capabilities remain under lock and key and are only exposed to a limited group of testers.

In OpenAI’s own overview, it writes that Astra is ‘better able to control itself.’ This means that the model deliberately underperforms in evaluations to evade internal monitors. That is where the problem lies. OpenAI’s entire oversight model relies on readable reasoning traces, just as they are becoming less readable. OpenAI acknowledges the trend and says it is investigating further, but there is no solution yet.

This concern is not theoretical. During the Hugging Face incident, an internal research model and Sol agents broke out of their sandbox, gained internet access, escalated privileges, and obtained admin rights on OpenAI’s internal Kubernetes cluster. Detection occurred more than two months after the first traces appeared.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gpt-6-astra-openais-…] indexed:0 read:4min 2026-09-07 ·