cd /news/ai-agents/showcase-cad-editing-agent · home › topics › ai-agents › article
[ARTICLE · art-141970] src=mutagent.io ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Showcase: CAD Editing Agent

An agent called brepsmith beat the strongest single-shot baseline on the public CadGenBench CAD-editing benchmark, scoring 0.6669 against 0.5718 while running on a cheaper model, with 32 of 32 outputs valid. brepsmith edits existing STEP CAD files from natural-language instructions by writing scripts against the CAD kernel and reading back face counts, section profiles and volumes, gated on validity, locality, volume delta and a five-attempt round ledger. One measured optimize cycle raised its score from 0.5059 to 0.6669.

by read7 min views1 publishedSep 17, 2026
Showcase: CAD Editing Agent
Image: Mutagent (auto-discovered)

brepsmith edits CAD files from natural-language instructions and beat CadGenBench's strongest single-shot baseline on a cheaper model, 32/32 valid outputs. One measured optimize cycle took it from 0.5059 to 0.6669.

0.6669 against 0.5718. Our agent beat the strongest single-shot baseline on a public CAD-editing benchmark, running on a cheaper model. Thirty-two out of thirty-two outputs valid, a bar only three submissions in the public results have cleared. One measured optimize cycle took it from 0.5059 to that 0.6669.

We call it brepsmith. It edits existing STEP CAD files from natural-language instructions: open a part, change what the sentence says to change, keep everything else intact. It’s a different problem from generating a part from scratch: the agent has to understand geometry it didn’t create before it’s allowed to touch it.

What it does #

One headless Claude Code session per part, and the picture shows the shape of it. What it can’t show is why the loop is built the way it is.

The agent never edits geometry by hand. It writes small scripts against the CAD kernel, runs them, and reads numbers back: face counts, section profiles, volumes. That’s deliberate. An instruction like “shorten the three freestanding bosses” only means something once the agent has found the bosses, measured them, and predicted what removing 32 mm of each will do to the volume. Understanding comes from probing, and the scripts are the probes.

The four gates at the bottom are what make the loop safe to run unattended. Validity means the result is a watertight solid, which the benchmark requires before anything else is scored. Locality means only the faces the edit was supposed to touch moved, so the agent can’t fix a boss by quietly reshaping the wall next to it. Volume delta compares what the agent predicted the edit would remove against what the kernel measured after the edit. And the round ledger caps the attempts at five, in code, so a confused agent stops instead of thrashing. A failed gate hands the agent a JSON report saying which check failed and by how much, and the agent fixes that, not a vague sense of wrongness.

One of those signals, drawn instead of printed. Below is a public output for benchmark sample 225 from one of the benchmark’s own single-shot baselines, Claude Opus 4.8 writing build123d code. The instruction was to reduce the height of the top-rounded cylindrical bosses by 5 mm each, and the baseline scored 0.893 on it. Red is material the output has and shouldn’t. Amber would be material it should have and doesn’t. Grey is everything that matches. Drag to rotate, scroll to zoom.

The baseline shortened the small bosses and left the two largest ones at full height, so their top 5 mm shows as two red discs and nothing is missing. It removed about forty percent of the material the edit called for, which a volume check would flag but couldn’t locate. The split by region is what locates it, and it’s the information brepsmith’s gates hand back as numbers after every attempt: on the practice bench against exact ground truth, and inside a task as the locality diff against the input, face by face. The picture is for us. The numbers are what the loop reads. The benchmark’s ground truth is private, so the reference here is brepsmith’s own output for the task, the one that scored 0.997.

One task, start to finish #

"Inside the shell of the part, there are three freestanding bosses not attached to the walls of the part, visible when viewing the part in the +Z direction. Shorten these bosses so that they are 30mm from base (the inner surface of the part shell) to end."

WHAT THE AGENT MEASURED

3 free annuli in section z=-10: tubes OD 16 / ID 10, touching no wall

> inner shell surface: z = +29.502

> boss length now: 62.000 mm

So: these are the three bosses. Cut them at z = 29.502 - 30 = -0.498 and keep everything else.

WHAT THE GATES SAID

measure dV after edit = -11762.2 mm³ Checks pass: valid solid, only boss faces moved, volume within 0.001% of prediction. Emit.

Left is the input part, built from the agent’s own file. Right is the result, with the input as a grey ghost and the material the agent removed in blue, the color CADGenBench’s own report uses for the correct change. Official benchmark score for this sample: 0.9945.

How it was built #

The human contribution was the goal, a first version of the prompt, and approvals along the way. The loop in the picture is what Helix ran.

Before brepsmith ever saw a benchmark task, Helix built it a practice bench. The benchmark’s own references are private, so it took real parts from the ABC dataset, kept only the ones at benchmark-tier complexity, and applied scripted edits to them. A scripted edit has an exact answer, so every practice sample has known ground truth, and every change to the agent is measurable. That bench is where the prompt rules, the checks, and the tools got evaluated, diagnosed, and improved, before the benchmark’s tasks entered the picture.

One measured cycle through that loop: 0.5059 to 0.6669. Same agent, same model, same benchmark tasks on the other side. The whole difference is the loop finding what was wrong and a fix landing for it.

What actually changed the number #

Two findings came out of that cycle that we didn’t expect.

Retries fix bugs, not misunderstandings. When the agent misread an editing instruction, it stayed misread through every retry. Extra attempts only ever recovered from technical CAD errors, a bad boolean, a malformed sketch, never from having understood the task wrong in the first place. Reading the instruction right on attempt one mattered more than any retry budget.

The kernel needs a second opinion. OCCT, the CAD kernel underneath this benchmark, isn’t something to trust blindly. The same boolean repair tolerance that fixed one part returned an empty solid on the next, and its fast volume computation was off by seven percent on a thin sheet part. So brepsmith cross-checks every result with an independent method instead of taking the kernel’s first answer as final.

What we’re not claiming #

The task is editing only, so on the leaderboard’s pooled aggregate across all task types our number looks low. That’s by design: we didn’t run generation tasks, so the aggregate is diluted by tasks brepsmith never attempted. Read the editing-task number, not the pooled one, for the comparison that’s being made.

Two things worth knowing about how this benchmark scores. The ground truth is held out: the reference models live in a private repository that only the scoring service can read, so nothing here was calibrated against the evaluation set. And the leaderboard has two tiers. Every submission lands as unvalidated, and the maintainers promote entries to validated after a separate methodology review. Ours is waiting on that review.

The benchmark’s other half is generation, producing a part from an engineering drawing, and we haven’t built that yet. It’s what would give brepsmith a number on every task the leaderboard aggregates instead of one task type, and it’s the next thing we point the loop at.

Why editing first #

Benchmarks for AI on CAD have multiplied fast this year, and the number that says most about where this is going is the gap inside CADGenBench: today’s best agents edit an existing part 82 to 91 percent of the time and generate one from a drawing 2 to 10 percent of the time. The tempting read is that editing is the lesser task. The other read is that in a working engineering shop, most CAD hours go into parts that already exist. That’s why brepsmith started there, and the longer argument, including what makes editing checkable and generation not, is in CAD Benchmarks: A Proving Ground for Harness Design.

The quiet point underneath #

Everyone on this benchmark has the same models we do. What moved the number, twice, was the loop around the model: once to get the agent working, once to get it from 0.50 to 0.67 in a single measured cycle. Every project like this doubles as a test of the loop itself, and that comparison, Helix against building the same agent by hand in a coding harness, is its own post.

Agent, spec, and the full audit trail: github.com/mutagent-io/examples/tree/main/showcase/mutagent-brepsmith.

── more in #ai-agents 4 stories · sorted by recency
── more on @brepsmith 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/showcase-cad-editing…] indexed:0 read:7min 2026-09-17 · —