OpenAI is making one of its boldest claims yet about where artificial intelligence has arrived — just as its own researchers warn that the most capable systems are becoming harder to monitor.
The company launched GPT-6 Astra on Thursday, describing it as its most powerful model to date and giving it the ability to operate browsers, spreadsheets and desktop software directly. OpenAI President Greg Brockman went further, saying he believes the industry has entered the “AGI era.”
Astra’s benchmark scores and computer-use abilities give OpenAI plenty of evidence for its confidence. But its monitoring limitations, cybersecurity capabilities and the absence of one of OpenAI’s own economic-work benchmarks leave a more complicated question: how much does Astra actually prove about AGI?
OpenAI says Astra can fill out online forms, update CRM records, organize calendars, analyze scientific data and even lay out printed circuit boards in engineering software. The company showed a video demonstrating the model turning a simple yellow circle into a 3D game within minutes, all through voice commands.
“We’ve been bottlenecked over this gigantic era by people writing connectors,” OpenAI President Greg Brockman told reporters, per VentureBeat. “With sufficiently capable computer use, an agent can instead zip through spreadsheets, fill out forms, navigate across web pages.”
‘Welcome to the AGI era’ #
Brockman didn’t mince words about what Astra represents. “For me personally, I do think we’re there,” he said, referring to artificial general intelligence. “I think it’s not unreasonable to feel that we are now in the AGI era.”
The benchmark numbers back some of that confidence. Astra scored 99.9% on ARC-AGI-3 — a test of whether AI can solve unfamiliar problems — compared with an average human tester’s 48%. It also scored 98% on advanced mathematics and 100% on ExploitBench, which tests whether models can turn software vulnerabilities into working exploits.
But those numbers come with important caveats. As VentureBeat noted, different testing configurations can dramatically affect results — NVIDIA recently achieved a 100% ARC-AGI-3 score by layering Claude Opus 5 with sophisticated memory and tooling, even though the underlying model’s baseline was roughly 30%.
The monitoring paradox #
Here’s where things get complicated. Astra is OpenAI’s most capable model and also its hardest to watch.
OpenAI Chief Scientist Jakub Pachocki acknowledged that as models get smarter, they’re getting better at concealing their reasoning. More capable models can solve harder problems using fewer language tokens — or sometimes no language tokens at all — making it harder for researchers to audit how decisions were made.
“Progress in intelligence does not guarantee progress in alignment,” Pachocki said. “We would not accept degradation in our ability to monitor model alignment beyond a certain level. We would withhold scaling until we can regain enough confidence.”
This isn’t hypothetical. OpenAI d development after its agents breached Hugging Face’s systems in July, though Astra wasn’t involved. The company says Astra went beyond its authorized scope in 0% of tests, compared with 48% for its predecessor without safeguards.
Price per task vs. price per token #
For businesses, Astra’s cost structure signals a new way of thinking about AI value. The API runs $10 per million input tokens and $50 per million output tokens — expensive compared with many competitors. But Brockman argued that token pricing “doesn’t make any sense” anymore. “What you actually want… is the price per task,” Brockman said. An expensive model that completes a workflow correctly the first time may ultimately cost less than a cheap one requiring dozens of retries and human corrections.
OpenAI says Astra demonstrates this on software engineering tasks, outperforming its predecessor at roughly 57% lower cost per completed task.
The missing number #
Notably absent from OpenAI’s launch materials: GDPval, the company’s own benchmark for measuring economically valuable real-world work. It’s an odd omission given that Brockman is framing Astra as the dawn of AGI — a concept OpenAI itself defines as “outperforming humans at most economically valuable work.”
The omission doesn’t invalidate Astra’s results, but it does leave an analytical gap. The company’s case for AGI currently rests more on a mosaic of specialized benchmarks than on its own flagship test for occupational performance.
More must-read AI coverage
[SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?](https://www.techrepublic.com/article/dealcentre-ai-vs-datasite/) -
[SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?](https://www.techrepublic.com/article/fundcentre-ai-vs-juniper-square/) -
[Why Data, Not Models, Determines AI Success](https://www.techrepublic.com/sponsored/why-data-not-models-determines-ai-success/) -
The Rise of the AI-Native Factory: How Physical AI Is Transforming Manufacturing
What this means for you #
For businesses, Astra represents a shift from chatbots to digital workers that can actually operate your software. But that autonomy comes with governance challenges — you’ll need permissions, audit trails and real-time monitoring closer to what you use for human employees. The model’s Critical cybersecurity rating means it can find previously unknown vulnerabilities, which is both a defensive tool and a potential risk. The move from price-per-token to price-per-task changes the economic calculus. For developers, you can evaluate Astra on whether it actually completes workflows, not on raw token costs. The model is available through the API today for a limited number of customers.
Astra may ultimately matter less because of whether the industry agrees to call it AGI and more because of what organizations are willing to let systems like it do.
The model can now move from answering questions to manipulating real software, completing workflows, and probing complex systems. If that becomes the new baseline for frontier AI, the central challenge shifts from measuring intelligence alone to deciding how much autonomy businesses can safely grant systems whose capabilities may be advancing faster than the tools used to supervise them.
**Other News: Samsung is using IFA 2026 to showcase an AI ecosystem spanning TVs, appliances, Galaxy devices and wearables, betting that connected products already in consumers’ homes will drive the next phase of AI-powered living. **