cd /news/ai-research/opus-5-5-scores-75-6-on-part-catalog… · home topics ai-research article
[ARTICLE · art-138606] src=partcatalogbench.adamjohnson.site ↗ pub= topic=ai-research verified=true sentiment=· neutral

Opus 5.5 Scores 75.6% on Part Catalog Bench

GPT-6 Astra scored 87.4% on the Part Catalog Bench, the highest of 13 models tested on 119 questions drawn from the 5,445-page Ford 1960–68 Car Master Parts and Accessories Catalog, according to benchmark creator results updated September 23, 2026. Claude Opus 5.5 placed second at 75.6% and Claude Fable 5.1 third at 61.3%, while the newly added GPT-6 Sol and GPT-6 Luna scored 55.5% and 21.0%; Opus 5.5 cost about $4.39 for the full question set versus $16.22 for Astra. The benchmark's author built it after AI tools struggled to interpret parts diagrams while he worked on a 1966 Ford Thunderbird, and all models were accessed through OpenRouter with one answer per question.

read4 min views1 publishedSep 23, 2026
Opus 5.5 Scores 75.6% on Part Catalog Bench
Image: source

I bought a 1966 Ford Thunderbird about a year ago, and I do most of the work on it myself. While working on the car, I often gave AI tools the manuals and parts catalogs and asked questions. They frequently had trouble understanding the diagrams.

I built this benchmark to find out what they could and couldn’t do, and to track whether they get better over time.

GPT-6 Astra has the highest score at 87.4%, followed by Claude Opus 5.5 at 75.6% and Claude Fable 5.1 at 61.3%. Opus costs about $4.39 for the full set of questions, compared with $16.22 for Astra.

Updated Sept 23, 2026

Me and my son in my 1966 Ford Thunderbird, the car that inspired me to create this benchmark. I used ChatGPT to remove a car and house in the background. The removal of my eyes was unsolicited.

Observed Pareto frontier 95% confidence intervalHigher and further left is better

Click a model to jump to its breakdown below.

Updates

September 23, 2026 — Three new models

I added Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna. They scored 75.6%, 55.5%, and 21.0%, respectively, on the same 119 questions with the same run settings. All three returned an answer to every question.

Opus 5.5 moves into second place behind GPT-6 Astra. The table, chart, downloads, and example answer pages now include all 13 models.

September 8, 2026 — First release

I published the benchmark with 119 questions across 10 exercises, and results for these ten models:

This is the front suspension diagram for the 1966–67 Ford Falcon and Ford Fairlane, with variations for the 1967 Ford Mustang. One exercise contains several questions about the same page. Here are three from this exercise, exactly as written in the benchmark.

Ford Car Master Parts and Accessories Catalog, illustration 30–18. Reproduced from the Forel edition. Click the image to open it at full size.

Easy · Part connections

Bolt 380571-S is used to attach part 3468 to what part?

Part 5495 passes through a series of other parts. List all the part numbers for these parts, in order from top to bottom as installed on the vehicle. Do not include 5495 itself.

Each model gets the image and one question at a time, along with instructions for reading the catalog and formatting its answer. It doesn’t see the other questions or their answers.

All the exercises use diagrams from the Ford 1960–68 Car Master Parts and Accessories Catalog, a 5,445-page reference Ford dealerships used to look up parts before computerized parts systems. I wrote the questions and answers myself.

To run the benchmark yourself, you’ll need your own PDF copy of the catalog. You can purchase it from Forel (product D10063).

What the model sees

A question, the relevant page images, and instructions on how to use the catalog. Images are rendered at 300 DPI, with any crops or highlights specified for that question. The PDF’s OCR text isn’t sent. None of the current questions requires looking something up in the text catalog.

Scoring

Every question has the same weight, regardless of difficulty. Part numbers, lists, and other structured answers are checked against the answer key. Free-text answers are graded against a rubric and can earn partial credit. Partial credit for incomplete lists is reported separately from the main score.

Run settings

These models were accessed through OpenRouter.

run settings…

Each score uses one answer per question. Some answers are reused from earlier runs when the question, image, and settings match. The assembly date is when those results were combined, not necessarily when every answer was generated.

Error bars

Several questions share each diagram, so the confidence intervals resample whole exercises rather than treating every question as independent. With only ten exercises, the intervals are still wide. They don’t capture how answers might change on a second run, or errors made by the grader.

Failed answers

A model gets zero if it runs out of output tokens or returns no final answer. Service errors can be retried. Models with unfinished runs or unresolved grading errors aren’t listed yet. A service failure doesn’t tell us how the model would have answered.

Cost and time

Cost is what the provider reported for the saved answers, including reused ones. Grading costs are listed separately. Paid retries that weren’t recorded may be missing, so these figures won’t necessarily match the bill. Response times also depend on provider load and how much reasoning the model does.

── more in #ai-research 4 stories · sorted by recency
── more on @gpt-6 astra 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/opus-5-5-scores-75-6…] indexed:0 read:4min 2026-09-23 ·