GPT-6 Astra Benchmarks Revealed: How OpenAI Says It Compares With Claude and GPT-5.6 OpenAI has unveiled GPT-6 Astra, reporting benchmark results that show significant gains over its predecessor GPT-5.6 Sol and several Claude models in coding, scientific reasoning, and professional tasks, though Anthropic's Claude Fable 5.1 retains the lead on Humanity's Last Exam with tools (65% vs. Astra's 57.2%). Astra scored 64.6% on Terminal-Bench Science 0.1, 57.9% on Terminal-Bench 4.0, 97.6% on FrontierMath Tier 4 (v2), 96% on GPQA Diamond, 41.4% on AutomationBench, and 95.9% on BenchCAD, with OpenAI highlighting its computer use capabilities and Critical cybersecurity threshold under its Preparedness Framework. The model is initially rolling out to a limited group of organizations, with access expanding to ChatGPT Plus, Pro, Business, and Enterprise users, as well as via the OpenAI API and Amazon Web Services. GPT-6 Astra Benchmarks Revealed: How OpenAI Says It Compares With Claude and GPT-5.6 GPT-6 Astra shows significant improvements in coding, scientific reasoning, and professional tasks, surpassing previous models OpenAI has unveiled GPT-6 Astra https://www.ibtimes.co.uk/openai-astra-ai-agi-progress-1816510 alongside benchmark results that suggest major gains in coding, scientific reasoning, computer use and professional tasks. The company describes Astra as its most capable and aligned model yet. Its published results place it ahead of GPT-5.6 Sol and several Claude models on many tests, although Claude Fable 5.1 retains the lead on at least one prominent evaluation. The figures were published by OpenAI, so they should be understood as company-reported results rather than a complete independent verdict on real-world performance. Benchmark comparisons can also depend on the reasoning setting, tool access and spending limits used for each model. The results therefore offer a useful snapshot of performance under OpenAI's test conditions, but they do not guarantee the same ranking for every workplace, coding or research task. How GPT-6 Astra Compares With Claude According to OpenAI's published benchmarks, Astra scored 64.6 per cent on Terminal-Bench Science 0.1, which measures whether AI agents can complete scientific workflows using code and terminal tools. Claude Fable 5.1, developed by Anthropic https://www.ibtimes.co.uk/ai-driven-protein-design-claude-breakthrough-1815432 , scored 52.6 per cent in OpenAI's comparison, while GPT-5.6 Sol achieved 22.4 per cent. Astra also scored 57.9 per cent on Terminal-Bench 4.0, which covers terminal-based work including software engineering, system configuration and data analysis. Claude Fable 5.1 reached 55.8 per cent, while GPT-5.6 Sol scored 37.3 per cent. However, Astra did not lead every assessment. Claude Fable 5.1 scored 65 per cent on Humanity's Last Exam with tools, compared with Astra's 57.2 per cent. OpenAI did not publish a GPT-5.6 Sol result for that test in the same table. Astra Shows Gains Over GPT-5.6 Sol Some of Astra's largest reported improvements appeared in scientific and professional evaluations. On FrontierMath Tier 4 v2 , Astra scored 97.6 per cent, ahead of Claude Fable 5.1 at 87.8 per cent and GPT-5.6 Sol at 83 per cent. Astra recorded 96 per cent on GPQA Diamond, compared with 94.6 per cent for GPT-5.6 Sol and 93.7 per cent for Claude Fable 5.1. The model also scored 41.4 per cent on AutomationBench, which evaluates professional computer tasks. Claude Fable 5.1 achieved 31.4 per cent, while GPT-5.6 Sol reached 18.1 per cent. On BenchCAD, Astra recorded 95.9 per cent, compared with 84.3 per cent for Claude Fable 5.1 and 83.3 per cent for GPT-5.6 Sol. OpenAI Highlights Computer Use and Safety OpenAI says Astra can handle complex, multi-step work involving browsers, software tools, documents, spreadsheets and presentations. The company's system card says Astra is its first model to reach the Critical cybersecurity capability threshold under its Preparedness Framework. OpenAI notes Astra demonstrated advanced abilities to identify previously unknown security vulnerabilities under controlled evaluation. The company said it strengthened safeguards against misuse and unauthorised actions before release https://www.ibtimes.co.uk/openai-gpt6-astra-task-automation-cybersecurity-1817833 . The system card also identifies limitations, including reduced visibility into parts of Astra's internal reasoning compared with GPT-5.6 Sol. When Will GPT-6 Astra Be Available? GPT-6 Astra is initially rolling out to a limited group of organisations. OpenAI says access will expand over the following days to ChatGPT Plus, Pro, Business and Enterprise users, as well as through the OpenAI API and Amazon Web Services. The benchmark results indicate a substantial advance over GPT-5.6 Sol across several categories. Still, performance varies by task, and OpenAI's own results show that Astra does not outperform Claude on every test. Real-world comparisons will ultimately depend on factors including accuracy, speed, cost, reliability and how each model performs outside controlled benchmark conditions. © Copyright IBTimes 2026. All rights reserved.