# Turns out, Astra cannot build manufacturing CAD yet

> Source: <https://interpretai.tech/benchmark/cad>
> Published: 2026-09-18 06:47:29+00:00

Frontier models asked to design production-grade CAD assemblies, graded on geometry, editability and manufacturability. Multiple task families, multiple graded tasks.

Snapshot of September 16, 2026

| Task family | claude fable 5.1claude code · max | deepseek v4 flash vision expopencode · max | gemini 3.8 flashantigravity · high | gpt 6 astracodex · max | grok 4.6grok build · xhigh | 
|---|---|---|---|---|---|
| Task Family A | 55% ± 0% | – | 24% ± 3% | 40% ± 15% | 19% ± 2% | 
| Task Family B | 52% ± 22% | 20% ± 0% | 17% ± 3% | 74% ± 8% | 13% ± 5% | 
| Task Family C | 55% ± 0% | 18% ± 4% | 19% ± 5% | 39% ± 2% | 18% ± 3% | 
| Task Family D | 36% ± 13% | 30% ± 0% | 17% ± 8% | 52% ± 5% | 22% ± 8% | 

The families shown are a sample of the benchmark's task families. A cell is the mean of every graded rollout the model completed in that task family, with the sample standard deviation over those rollouts. A dash is a model with no completed rollout in that family: its runs failed to finish. A score of 60% or above is a pass.

The x axis prices those tokens at each provider's published list rate as of September 2026.
