Turns out, Astra cannot build manufacturing CAD yet A benchmark of frontier models on production-grade CAD assembly design found none reached the 60% pass threshold as of September 16, 2026, with GPT 6 (Astra, Codex at max) scoring highest at 74% ± 8% on Task Family B but only 40% ± 15% on Task Family A and 52% ± 5% on Task Family D. Claude Fable 5.1 (Claude Code at max) scored 55% ± 0% on Task Families A and C, 52% ± 22% on B and 36% ± 13% on D, while Gemini 3.8 Flash (Antigravity at high) topped out at 24% ± 3% and Grok 4.6 (Grok Build at xhigh) at 22% ± 8%. DeepSeek V4 Flash Vision (OpenCode at max) completed only Task Families B, C and D, scoring 20% ± 0%, 18% ± 4% and 30% ± 0% respectively, with dashes marking models whose runs failed to finish. Frontier models asked to design production-grade CAD assemblies, graded on geometry, editability and manufacturability. Multiple task families, multiple graded tasks. Snapshot of September 16, 2026 | Task family | claude fable 5.1claude code · max | deepseek v4 flash vision expopencode · max | gemini 3.8 flashantigravity · high | gpt 6 astracodex · max | grok 4.6grok build · xhigh | |---|---|---|---|---|---| | Task Family A | 55% ± 0% | – | 24% ± 3% | 40% ± 15% | 19% ± 2% | | Task Family B | 52% ± 22% | 20% ± 0% | 17% ± 3% | 74% ± 8% | 13% ± 5% | | Task Family C | 55% ± 0% | 18% ± 4% | 19% ± 5% | 39% ± 2% | 18% ± 3% | | Task Family D | 36% ± 13% | 30% ± 0% | 17% ± 8% | 52% ± 5% | 22% ± 8% | The families shown are a sample of the benchmark's task families. A cell is the mean of every graded rollout the model completed in that task family, with the sample standard deviation over those rollouts. A dash is a model with no completed rollout in that family: its runs failed to finish. A score of 60% or above is a pass. The x axis prices those tokens at each provider's published list rate as of September 2026.