# I'm afraid of spiders. So I made AI look at 2k of them

> Source: <https://labqoat.com/blog/how-well-can-ai-identify-spiders>
> Published: 2026-09-21 09:00:37+00:00

[← All posts](https://labqoat.com/blog)

# I’m Afraid of Spiders. So I Made AI Look at 2,000 of Them.

## In this article

I’m afraid of spiders. A lot. 🕷️

My identification system currently consists of “the one with long legs” and “the short but big one.” There’s also “where the fuck did it go,” but that’s more of an emergency than a classification.

So I got curious: how well could AI actually identify them?

I tested nine models on the same 2,000 spider photos to see how often they got the species right.

(Looking at this many spiders was not a comfortable experience.)

This post was heavily inspired by [Piotr Migdał’s post on identifying mushrooms with AI](https://quesma.com/blog/mushroom-llm-vision/).

## [First, approximately ten seconds of biology](#first-approximately-ten-seconds-of-biology)

A **family** is a broad group of related spiders. A **genus** is a smaller group inside it, and a **species** is the specific kind of spider.

For example: **Araneidae → Araneus → Araneus diadematus**, the [European garden spider](https://en.wikipedia.org/wiki/Araneus_diadematus). In that last name, *Araneus* is the genus.

That’s enough biology for now.

## [How the benchmark works](#how-the-benchmark-works)

Each model got the same **2,000 photos**, covering **671 species and subspecies**, and a list of **20 possible names** for each photo, including the correct answer.

The task was to identify the spider’s species from the photo and return the matching name from the list.

I gave each model the same possible answers so I could compare how well they distinguished the species, with fewer ambiguities in scoring. That makes this a **multiple-choice identification test**, and the choices themselves can help.

I built the dataset from research-grade [iNaturalist](https://www.inaturalist.org/) observations, using the community’s species identifications as the expected answers. The species list comes from a Polish spider checklist, but the photos were taken worldwide. After filtering out label mismatches and unsuitable images, 2,183 photos remained, and I used the same [2,000-photo subset](https://github.com/qforge-dev/spider-bench/blob/main/data/benchmarks/runs/gemini-20260910-230024/tasks.jsonl) for every model.

## [Let’s look at a few examples](#lets-look-at-a-few-examples)

I’d call any of these “spider,” but… 😅

According to [iNaturalist](https://www.inaturalist.org/observations/24731930), it’s a **Red-bellied Jumping Spider** (*Philaeus chrysops*).

### What did the models say?

9 of 9 picked the expected species.

| Model answers for Philaeus chrysops |  |  | 
|---|---|---|
| Model | Answer | Result | 
|---|---|---|
| Gemini 3.8 Flash | Philaeus chrysopsCorrect | Correct | 
| GPT-6 Astra | Philaeus chrysopsCorrect | Correct | 
| Claude Fable 5.1 | Philaeus chrysopsCorrect | Correct | 
| Muse Spark 1.3 <sup>*</sup> | Philaeus chrysopsCorrect | Correct | 
| GLM 5.3 Flash | Philaeus chrysopsCorrect | Correct | 
| GPT-5.6 Sol | Philaeus chrysopsCorrect | Correct | 
| DeepSeek V4.1 Flash | Philaeus chrysopsCorrect | Correct | 
| GPT-5.6 Terra | Philaeus chrysopsCorrect | Correct | 
| GPT-5.6 Luna | Philaeus chrysopsCorrect | Correct | 

Everyone got this one right.

## See the 20 choices the models received

1. 1.Nigma walckenaeri
2. 2.Zelotes aeneus
3. 3.Attulus inexpectus
4. 4.Philaeus chrysops
5. 5.Gnaphosa lucifuga
6. 6.Savignia frontata
7. 7.Sittisax saxicola
8. 8.Kishidaia conspicua
9. 9.Evarcha laetabunda
10. 10.Metellina mengei
11. 11.Attulus rupicola
12. 12.Araniella displicata
13. 13.Leptorchestes berolinensis
14. 14.Neon valentulus
15. 15.Attulus terebratus
16. 16.Marpissa pomatia
17. 17.Cybaeus tetricus
18. 18.Nuctenea umbratica
19. 19.Euophrys frontalis
20. 20.Larinioides ixobolus

<sup>*</sup> Muse Spark 1.3 uses the Contributor tier.

Selected examples; overall scores use all 2,000 photos.

## [Okay, who knew the spiders?](#okay-who-knew-the-spiders)

**Gemini 3.8 Flash got 997 out of 2,000 correct: 49.85%**, the highest score in these runs.

GPT-6 Astra followed at **47.50%**, then Claude Fable 5.1 at **43.10%** and Muse Spark 1.3 at **39.10%**. The top two were separated by 47 photos. I’d read that as a result on this set of spiders, not a universal ranking of how much the models know about them.

Same 2,000 photos for every model. Wrong and failed answers stay in the denominator.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

## See the exact numbers

| Exact-species accuracy. Each run contains 2,000 photos. |  |  | 
|---|---|---|
| Model | Correct | Accuracy | 
|---|---|---|
| Gemini 3.8 Flash | 997 | 49.85% | 
| GPT-6 Astra | 950 | 47.50% | 
| Claude Fable 5.1 | 862 | 43.10% | 
| Muse Spark 1.3 <sup>*</sup> | 782 | 39.10% | 
| GLM 5.3 Flash | 735 | 36.75% | 
| GPT-5.6 Sol | 673 | 33.65% | 
| DeepSeek V4.1 Flash | 528 | 26.40% | 
| GPT-5.6 Terra | 456 | 22.80% | 
| GPT-5.6 Luna | 430 | 21.50% | 

## [How wrong is wrong?](#how-wrong-is-wrong)

Gemini 3.8 Flash missed the exact species in 1,003 photos. But in **871 of those cases**, it named another spider from the same family. That’s almost **87% of its misses**. In 276 cases, it even got the genus right.

GPT-6 Astra and Claude Fable 5.1 showed a similar pattern. Their exact-species scores were **47.50% and 43.10%**, but their answers belonged to the correct family **92.00% and 91.45%** of the time. Most of their mistakes happened between species within the same family.

Muse Spark 1.3 identified more exact species than GLM 5.3 Flash. GLM’s answers, however, belonged to the correct family slightly more often: **89.40% versus 88.70%**.

That makes the errors more interesting than a simple correct-or-wrong score suggests. There’s a substantial difference between getting the broad group right and identifying the particular species.

- Correct species
- Same genus only
- Same family only
- Wrong family
- Failed response

Each answer appears in one segment. “Only” means the more specific identification was wrong. Failed responses include invalid names, refusals, empty answers, truncations, and request errors.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

## See the exact numbers

| How close were the answers?. Each run contains 2,000 photos. |  |  |  |  |  | 
|---|---|---|---|---|---|
| Model | Correct species | Same genus only | Same family only | Wrong family | Failed response | 
|---|---|---|---|---|---|
| Gemini 3.8 Flash | 997 | 276 | 595 | 128 | 4 | 
| GPT-6 Astra | 950 | 312 | 578 | 156 | 4 | 
| Claude Fable 5.1 | 862 | 347 | 620 | 169 | 2 | 
| Muse Spark 1.3 <sup>*</sup> | 782 | 306 | 686 | 226 | 0 | 
| GLM 5.3 Flash | 735 | 322 | 731 | 209 | 3 | 
| GPT-5.6 Sol | 673 | 264 | 684 | 378 | 1 | 
| DeepSeek V4.1 Flash | 528 | 328 | 710 | 342 | 92 | 
| GPT-5.6 Terra | 456 | 272 | 718 | 554 | 0 | 
| GPT-5.6 Luna | 430 | 248 | 671 | 606 | 45 | 

Counts, not cumulative percentages. Each row adds up to 2,000.

## [Accuracy versus cost](#accuracy-versus-cost)

Muse Spark 1.3 was the cheapest model here and still finished fourth, ahead of five more expensive models. At its Contributor pricing, the entire run cost **$0.93 for 782 correct identifications**.

Moving to Gemini 3.8 Flash brought that total to 997. That’s **215 more correct answers for another $26.18**, an improvement of **10.75 percentage points** at roughly **29 times the cost**. 💸

Spending beyond that didn’t improve the results. Claude Fable 5.1 cost about twice as much as Gemini, and GPT-6 Astra cost **3.6 times as much**. Both scored lower. Gemini delivered the highest accuracy here; Muse offered a cheaper trade-off if identifying fewer species was acceptable.

Estimated totals for 2,000 photos. The dollar axis is logarithmic: spacing represents cost ratios.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

## See the exact numbers

| Accuracy versus estimated cost. Each run contains 2,000 photos. |  |  |  | 
|---|---|---|---|
| Model | Correct | Accuracy | Cost | 
|---|---|---|---|
| Gemini 3.8 Flash | 997 | 49.85% | $27.11 | 
| GPT-6 Astra | 950 | 47.50% | $97.53 | 
| Claude Fable 5.1 | 862 | 43.10% | $54.19 | 
| Muse Spark 1.3 <sup>*</sup> | 782 | 39.10% | $0.93 | 
| GLM 5.3 Flash | 735 | 36.75% | $1.46 | 
| GPT-5.6 Sol | 673 | 33.65% | $42.89 | 
| DeepSeek V4.1 Flash | 528 | 26.40% | $8.52 | 
| GPT-5.6 Terra | 456 | 22.80% | $20.53 | 
| GPT-5.6 Luna | 430 | 21.50% | $4.99 | 

## [Is getting half wrong actually bad?](#is-getting-half-wrong-actually-bad)

Getting almost half right is respectable when the task is telling similar species apart. The gap between recognising a family and naming the exact species is where this gets difficult.

Try telling  apart. For *Araniella opisthographa*, [NatureSpot’s identification guidance](https://www.naturespot.org/species/araniella-opisthographa) calls for examination at high magnification. That’s the sort of detail a photo can leave out.

The code and full benchmark results are on [GitHub](https://github.com/qforge-dev/spider-bench).

*PS: If you actually know spiders, I’d love to hear what you think on [twitter](https://x.com/breeg554). How obvious are these mistakes to someone who knows what they’re looking at?*

P.P.S. Thankfully, I have a cat in charge of spider security at home.

[← All posts](https://labqoat.com/blog)Labqoat
