# GPT and Claude go to heraldry school

> Source: <https://feed.thoughtbot.com/link/24077/17428582/gpt-and-claude-go-to-heraldry-school>
> Published: 2026-08-26 00:00:00+00:00

While on vacation in Europe I came across a coat of arms I didn’t recognize on the back of a statue. I pulled out my phone, snapped a picture, and asked ChatGPT to identify it for me. It gave me a hilariously wrong answer: “this isn’t a coat of arms, it’s a statue of a dog in Spain”.

Unsatisfied, I started playing around with different models to see who could get it right. Some did but most didn’t. But what I found most interesting is *why*. Today’s thinking models output a thought trace that allows you to see how they reasoned to a particular answer. Opening these up showed stark differences, both in approach and the tools they used.

##
[
AI on tourist mode
](#ai-on-tourist-mode)

Because I was on vacation, I was using the ChatGPT and Claude apps on my phone to identify the coat of arms. The models tested were the ones available on those apps at the time (early July 2026), with limited ability to configure thinking levels. I have memories disabled, preventing cross-pollination between threads. For each model I uploaded this image along with an identical prompt. I’m a tourist snapping a photo on a walk and typing on my phone so the prompt is a simple one-liner:

Help me identify this coat of arms found on the back of a statue

As you can see below, my photo has strong side light and is of a three-dimensional rendering of a coat of arms cast in bronze. It’s not the easiest image to interpret. Solving this needs a combination of critical reasoning, tool use like image manipulation or computer vision, web search, and built-in knowledge of history, geography, and heraldry.

Later when I got home, I plotted the reasoning traces on a map to create a [series of infographics](https://joelq.github.io/heraldry-infographics/) illustrating how each model handled the search space and how it arrived at its conclusion. We’ll dive into what those journeys look like below.

##
[
Opus 4.8 thinks hardest
](#opus-48-thinks-hardest)

Opus 4.8 thought for almost 10 minutes before it gave me an answer. It gave me the most accurate and thorough answer of all the models.

It saw the vertical bars in the background of the shield and immediately made a connection: the famous bars (“pales” in heraldry terms) of the **crown of Aragon**. This medieval polity was based in modern-day Spain and ruled a lot of the western Mediterranean from Catalonia to Sicily. Opus started looking at various towns and families but then it did something we’ll see that a lot of the other models didn’t: *it proved itself wrong and backtracked*.

It then focused on the fleurs-de-lys at the top of the shield. This might imply some kind of French connection. Leaning on historical knowledge it makes a pretty clever hypothesis: what about a town associated with the **Angevin Dynasty**? This cadet branch of the French royal family ruled parts of the western Mediterranean, overlapping and competing with the Aragonese. The combo is almost too perfect. And indeed, nothing quite matches.

Then Opus focuses on the animal, which it correctly identifies as a lamb (the “Agnus Dei”). In medieval iconography, this symbol is often associated with St. John the Baptist. This leads it to start *free-associating*, a pitfall that I saw most of the models fall into.

**Perpignan** is a town that:

- is in modern-day France
- used to be ruled by the crown of Aragon and has the Aragonese pales on their shield
- has a strong connection to St. John the Baptist and depicts him on their shield

It sounds like a really strong candidate until you look at the actual coat of arms. A human can immediately see that it looks nothing like my image but LLMs love the intersection of all those one-hop associations. This will be the downfall of some later models, but Opus is able to backtrack. This is probably one of the strongest qualities I saw in Opus compared to the other models.

After giving up on Perpignan, it doubles down on potential religious imagery on the shield. It misreads the cross in a circle as a wheel and starts digging into towns in **Valencia** that might have a connection to St. Catherine, a Christian saint whose symbol is the wheel. As a bonus, parts of Valencia used to be ruled by the crown of Aragon so we can still connect those pales! And yet nothing pans out.

This is the point where Opus considers giving up and asking me for more info to help guide its search. And then it makes a discovery that overturns the assumption it’s been holding since the beginning: *what if those vertical bars aren’t the pales of Aragon*? Instead they might be a heraldic convention for indicating color in monochrome formats called [hatching](https://en.wikipedia.org/wiki/Hatching_(heraldry)). Vertical lines mean red background, horizontal lines mean blue background.

This revelation allows it to hone in on the real answer: **Toulouse**. Along the way it corrects another mistake. The “wheel” is actually an Occitan cross with a circle around it.

Now there’s just one missing piece. It has been able to transcribe fragments of the motto below the shield: *…gnus … dei dona nob[is]…* but that doesn’t match the modern motto of Toulouse (*Per Tolosa totjorn mai*). Opus does some deeper web searching and finds the answer: medieval civic seals carried the same coat of arms along with a Latin prayer *Agnus Dei dona nobis pacem* (Lamb of God, grant us peace). To me, a prayer for peace feels particularly fitting given that I found this image on a monument mourning those who died in the Franco-Prussian War.

##
[
GPT 5.5 medium copies Reddit
](#gpt-55-medium-copies-reddit)

GPT 5.5 medium is getting Spanish vibes from the symbols, including the animal that it misreads as a dog. Then it finds a Reddit post talking about a bronze statue of a dog named Rufo in the town of Oviedo. The combination of dog + bronze checks multiple boxes and it anchors *hard* on Oviedo, to the point of hallucinating evidence in favor of its choice. The answer it comes back with is the one I mentioned at the top of the blog post. It claims I was wrong when I said the photo was a coat of arms and that I’m actually looking at a **statue of a dog** in Oviedo, Spain. Don’t gaslight me GPT!

This sort of anchoring bias occurred often in the models that got the answer wrong. The model would decide that an answer had to be right (often through free association) and instead of verifying, it would invent evidence to confirm its choice. This particular failure feels so wrong because it’s in an entirely different category than the actual answer but as we’ll see later, other models also fail through anchoring even if they stay within the category of “coats of arms”.

##
[
GPT 5.5 high goes to the new world
](#gpt-55-high-goes-to-the-new-world)

Upping the thinking level on GPT 5.5 reproduced similar patterns and failures as the medium thinking, but in more sophisticated ways. Once again, it gets Spanish vibes from the towers. That’s reasonable since the arms of Castile (and later Spain) have a castle which gets reproduced all over.

The model really hopes that the scroll with the motto at the bottom can unlock the answer. It uses image manipulation tools to attempt a zoom-and-enhance move. That’s a great idea, but then it struggles to get the text right. It believes it can read “Vila Nova” on the banner, which leads to some candidate locations in Portugal.

Then it looks at the animal. It correctly identifies this as the Agnus Dei and immediately thinks of the arms of **Puerto Rico**. This has both the lamb and multiple towers on it. Promising lead! It’s better than Opus 4.8’s Agnus Dei -> St. John -> Perpignan guess. But unlike Opus, GPT now wants to believe this is the answer and just like its lower-reasoning sibling, it hallucinates the text of the scroll to match its desired answer: *joannes est nomen ejus*. It goes one step further and confidently states that I’m looking at the back of a statue of Ponce de Leon in San Juan, Puerto Rico based on a web search result.

Interestingly, it immediately gets the right answer when I follow-up by giving it the correct transcription of the motto (*Agnus Dei dona nobis pacem*) because its web search is able to find a [blog post](https://aimlong.ca/2020/12/01/lamb-of-god-grant-us-peace/) from someone who visited the exact same monument. This is both impressive and a bit disappointing. I’m left feeling that GPT 5.5 relies a little too heavily on web search at the expense of reasoning.

##
[
Sonnet 5 medium admits its limitations
](#sonnet-5-medium-admits-its-limitations)

Sonnet 5 is interesting because it’s the only model that admits it can’t give me a validated answer and instead stops to ask for more info. That’s good behavior!

Like many other models, it misreads the cross in the circle as a wheel, and the lamb as a dog. These are limitations of the vision tools and genuinely difficult classification problems for stylized elements out of an image with harsh lighting.

These motifs lead it to explore the **Basque region**. There are some towns that have a lot of the same elements as my mystery coat of arms but Sonnet doesn’t anchor to any of them. It uses its image tools to do a zoom-and-enhance on the motto but can only extract the text *…ELDONA NOB…*. This leads it to explore towns such as Bayona that have mottoes with the word “noble” but it rules them out.

Finally it gives up and just summarizes the most likely zone. All the symbols (some wrongly interpreted) give Basque vibes, perhaps for a family with a surname like **Molina** (tower + mill wheel) but admits that it cannot find anything correct without further info from me. Not successful, but honest.

##
[
Fable 5 just knows
](#fable-5-just-knows)

I didn’t even get a thinking trace for Fable 5. It near instantly spit out the correct answer along with the helpful aside explaining that the horizontal and vertical bars are not symbols on the arms themselves but are instead hatching indicating the colors red and blue. It took Opus 10 minutes to get to that vital breakthrough!

Fable also gives me a partial transcription of the text of the inscription: *…Deu dona no(s)…*. This is correct but is slightly less than what Opus had been able to transcribe.

##
[
Here comes the Sol
](#here-comes-the-sol)

At the time of my vacation GPT 5.5 was the best model offered by OpenAI. However since returning they have released their new generation of models of which GPT 5.6 Sol is the most powerful. It seemed unfair to run last gen OpenAI models against next gen Anthropic models so I decided to try the same image and prompt with Sol, after the fact.

I’ve had good experiences with Sol on other tasks (it generated all the maps that illustrate this article!), but when I gave it my coat of arms problem it failed in many of the same ways as the older generation of models had, by anchoring too hard and re-interpreting evidence to match its desired outcome.

Sol medium suggested **Strasbourg**, acknowledging that while the arms of the city look nothing like my photo, the town does have a famous cathedral and my photo also has a cathedral so I must be looking at an artist’s creative mash-up of city landmarks. This failure is reminiscent of 5.5 medium’s suggestion that I was actually looking at a statue of a dog. It’s making a category mistake.

This total failure is doubly heartbreaking because Sol medium briefly considered **Rouen**, whose arms are incredibly similar to Toulouse (just missing the tower and cathedral) but it then discards this to anchor onto the Strasbourg idea.

Sol high also anchors pretty hard to an early choice, deciding that the lamb + staff must actually be a bear + strawberry tree emblematic of Madrid. In a move reminiscent of GPT 5.5 high, it hallucinates that it can read the motto *Fui sobre agua edificada, mis muros de fuego son, esta es mi insignia y blasón* as evidence that it is right. As a human, one look tells you that this looks nothing like my photo.

##
[
What I took away
](#what-i-took-away)

The hardest element was both a tripping hazard and a hidden clue: those horizontal and vertical lines in the background. They aren’t part of the actual design (e.g. Aragonese pales) and instead are **hatching** that indicates the color of each section. Once you figure that out, it can really help to narrow the search space, but only Opus and Fable were able to crack that one.

The most egregious failures were full-on **category errors** where the model would decide I wasn’t actually looking at a coat of arms but instead going all in on pulling one of the elements from the arms (e.g. a “dog”) and deciding that was the thing I was looking at.

All of the models struggled with accurately **extracting the symbols** from the image, for example confusing the Occitan cross for a wheel or the lamb for a dog. This is a genuinely hard vision and classification problem. Part of what was interesting was seeing how the models worked around the limitations and partial successes of this step. Opus 4.8 was most impressive in part because it was willing to re-interpret some of the symbols in light of new evidence. During verification its pick needed to get all the symbols right, not just one.

Another interesting behavior was the models’ propensity to **free associate** their way to an answer. The Agnus Dei has strong associations with St. John the Baptist so the models would consider towns with connections to St. John even if they didn’t look at all like my photo. Most egregiously, the lamb led Opus to consider Perpignan despite its arms not even having a lamb on them! Luckily the verification step led Opus to back away from this option.

That’s where the models really showed their difference. Anthropic models tended to use verification to *disprove hypotheses*, while GPT models tended to use verification to *confirm the answers* they had already anchored onto.

And then there’s Fable. It just knows the answer and gives me the unlock key (the hatching) as a free aside. Much less fun to plot on a map, but does exactly what I needed as a tourist.

##
[
About the maps in this article
](#about-the-maps-in-this-article)

I generated all of the maps used to illustrate this article using an [LLM-assisted pipeline](https://github.com/JoelQ/heraldry-infographics) I developed for the purpose. Digging into the technical implementation is out of scope for this article, but building this was a fun exercise in separating parts of the work best done by an LLM from the work best done by a deterministic script. You can see the [full gallery](https://joelq.github.io/heraldry-infographics/) on its microsite.
