Astra's headline 95.9% score came only when the model could test and retry its own designs #
Videos of OpenAI's new GPT-6 Astra model designing detailed mechanical parts on command have raced across social media since its 3 September launch, yet engineers testing the tool warn that almost none of those parts are safe to actually manufacture.
The clips are striking. Astra has been filmed building a rotating turbocharger split into working parts, a camera broken into more than a hundred components, and a car model separated into several hundred individually named pieces, all generated from a single prompt. One robotics designer used it to lay out a full competition robot intake in about seven hours. Almost all of these came from independent creators, not from OpenAI's own launch materials.
The Score That Set off the Hype #
OpenAI released Astra on 3 September, calling it its most capable model yet and pointing to gains in computer use, coding, science, and computer-aided design (CAD). The figure that caught design professionals' attention was BenchCAD, a test that measures how closely an AI-built part matches a target shape.
Astra scored 95.9%, against 83.3% for the earlier GPT-5.6 Sol and 84.3% for Anthropic's Claude Fable 5.1, though OpenAI noted the Claude comparison used modified evaluation settings.
The Catch in the Fine Print #
The headline number carries a condition that rarely makes it into viral clips. Astra hit 95.9% only in an agentic setting, where it could render the part, measure it, and try again before submitting an answer, according to the BenchCAD project.
In the plain setting with no tools, Astra has published no score, and GPT-5.6 Sol still leads. A model that can check its own work is a real step forward, but it is a different claim from saying AI can now do engineering.
Why Engineers Aren't Sold #
The gap shows up fast in testing. The BenchCAD researchers recorded one part that scored 96.1% geometric similarity yet was still the wrong part, because a later feature quietly closed an opening the design was meant to keep.
A separate 2026 benchmark, neuralCAD-Edit, measured frontier models against ten professional designers on real edits and found the best model finished 53 percentage points behind the humans.
Robotics builders reviewing Astra's output on public forums were blunter, with several saying the results looked impressive but would take more work to fix than to draw correctly from scratch. Others questioned the running cost and whether handing design to a model teaches anyone anything.
What Astra Actually Changes #
Most experts expect Astra to sit above CAD software rather than replace it, reading instructions and writing code while the CAD system holds the exact geometry, dimensions, and constraints. That still matters for anyone whose job is downstream. In safety-critical work like aerospace or medical devices, a part that passes every check inside the software can still fail inspection or in service, and no benchmark catches that.
Cost is a live question too. OpenAI lists Astra at $10 (about £7.40) per million input tokens and $50 (about £37) per million output, and one tester estimated a single seven-hour design run burned through a few hundred dollars in usage.
For now, the safest use is the oldest rule in engineering. A qualified human still signs off the final drawing. © Copyright IBTimes 2026. All rights reserved.