cd /news/computer-vision/i-m-afraid-of-spiders-so-i-made-ai-l… · home topics computer-vision article
[ARTICLE · art-135713] src=labqoat.com ↗ pub= topic=computer-vision verified=true sentiment=· neutral

I'm afraid of spiders. So I made AI look at 2k of them

Gemini 3.8 Flash identified spider species correctly on 997 of 2,000 photos, a 49.85% accuracy rate that was the highest among nine AI models tested in a benchmark built from research-grade iNaturalist observations covering 671 species and subspecies. The multiple-choice test, run on the same 2,000-photo subset drawn from 2,183 filtered images, gave each model a list of 20 possible species names per photo including the correct answer. The benchmark was inspired by Piotr Migdał's post on identifying mushrooms with AI.

by read8 min views1 publishedSep 21, 2026
I'm afraid of spiders. So I made AI look at 2k of them
Image: source

← All posts

In this article #

I’m afraid of spiders. A lot. 🕷️

My identification system currently consists of “the one with long legs” and “the short but big one.” There’s also “where the fuck did it go,” but that’s more of an emergency than a classification.

So I got curious: how well could AI actually identify them?

I tested nine models on the same 2,000 spider photos to see how often they got the species right.

(Looking at this many spiders was not a comfortable experience.)

This post was heavily inspired by [Piotr Migdał’s post on identifying mushrooms with AI](https://quesma.com/blog/mushroom-llm-vision/).

## [First, approximately ten seconds of biology](#first-approximately-ten-seconds-of-biology)

A family is a broad group of related spiders. A genus is a smaller group inside it, and a species is the specific kind of spider.

For example: Araneidae → Araneus → Araneus diadematus, the European garden spider. In that last name, Araneus is the genus. That’s enough biology for now.

How the benchmark works #

Each model got the same 2,000 photos, covering 671 species and subspecies, and a list of 20 possible names for each photo, including the correct answer.

The task was to identify the spider’s species from the photo and return the matching name from the list.

I gave each model the same possible answers so I could compare how well they distinguished the species, with fewer ambiguities in scoring. That makes this a multiple-choice identification test, and the choices themselves can help.

I built the dataset from research-grade iNaturalist observations, using the community’s species identifications as the expected answers. The species list comes from a Polish spider checklist, but the photos were taken worldwide. After filtering out label mismatches and unsuitable images, 2,183 photos remained, and I used the same 2,000-photo subset for every model.

Let’s look at a few examples #

I’d call any of these “spider,” but… 😅

According to iNaturalist, it’s a Red-bellied Jumping Spider (Philaeus chrysops).

What did the models say?

9 of 9 picked the expected species.

Model answers for Philaeus chrysops
Model Answer Result
--- --- ---
Gemini 3.8 Flash Philaeus chrysopsCorrect Correct
GPT-6 Astra Philaeus chrysopsCorrect Correct
Claude Fable 5.1 Philaeus chrysopsCorrect Correct
Muse Spark 1.3 <sup>*</sup> Philaeus chrysopsCorrect Correct
GLM 5.3 Flash Philaeus chrysopsCorrect Correct
GPT-5.6 Sol Philaeus chrysopsCorrect Correct
DeepSeek V4.1 Flash Philaeus chrysopsCorrect Correct
GPT-5.6 Terra Philaeus chrysopsCorrect Correct
GPT-5.6 Luna Philaeus chrysopsCorrect Correct

Everyone got this one right.

See the 20 choices the models received #

  1. 1.Nigma walckenaeri
  2. 2.Zelotes aeneus
  3. 3.Attulus inexpectus
  4. 4.Philaeus chrysops
  5. 5.Gnaphosa lucifuga
  6. 6.Savignia frontata
  7. 7.Sittisax saxicola
  8. 8.Kishidaia conspicua
  9. 9.Evarcha laetabunda
  10. 10.Metellina mengei
  11. 11.Attulus rupicola
  12. 12.Araniella displicata
  13. 13.Leptorchestes berolinensis
  14. 14.Neon valentulus
  15. 15.Attulus terebratus
  16. 16.Marpissa pomatia
  17. 17.Cybaeus tetricus
  18. 18.Nuctenea umbratica
  19. 19.Euophrys frontalis
  20. 20.Larinioides ixobolus

<sup>*</sup> Muse Spark 1.3 uses the Contributor tier.

Selected examples; overall scores use all 2,000 photos.

Okay, who knew the spiders? #

Gemini 3.8 Flash got 997 out of 2,000 correct: 49.85%, the highest score in these runs.

GPT-6 Astra followed at 47.50%, then Claude Fable 5.1 at 43.10% and Muse Spark 1.3 at 39.10%. The top two were separated by 47 photos. I’d read that as a result on this set of spiders, not a universal ranking of how much the models know about them.

Same 2,000 photos for every model. Wrong and failed answers stay in the denominator.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

See the exact numbers #

Exact-species accuracy. Each run contains 2,000 photos.
Model Correct Accuracy
--- --- ---
Gemini 3.8 Flash 997 49.85%
GPT-6 Astra 950 47.50%
Claude Fable 5.1 862 43.10%
Muse Spark 1.3 <sup>*</sup> 782 39.10%
GLM 5.3 Flash 735 36.75%
GPT-5.6 Sol 673 33.65%
DeepSeek V4.1 Flash 528 26.40%
GPT-5.6 Terra 456 22.80%
GPT-5.6 Luna 430 21.50%

How wrong is wrong? #

Gemini 3.8 Flash missed the exact species in 1,003 photos. But in 871 of those cases, it named another spider from the same family. That’s almost 87% of its misses. In 276 cases, it even got the genus right.

GPT-6 Astra and Claude Fable 5.1 showed a similar pattern. Their exact-species scores were 47.50% and 43.10%, but their answers belonged to the correct family 92.00% and 91.45% of the time. Most of their mistakes happened between species within the same family.

Muse Spark 1.3 identified more exact species than GLM 5.3 Flash. GLM’s answers, however, belonged to the correct family slightly more often: 89.40% versus 88.70%.

That makes the errors more interesting than a simple correct-or-wrong score suggests. There’s a substantial difference between getting the broad group right and identifying the particular species.

  • Correct species
  • Same genus only
  • Same family only
  • Wrong family
  • Failed response

Each answer appears in one segment. “Only” means the more specific identification was wrong. Failed responses include invalid names, refusals, empty answers, truncations, and request errors.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

See the exact numbers #

How close were the answers?. Each run contains 2,000 photos.
Model Correct species Same genus only Same family only Wrong family Failed response
--- --- --- --- --- ---
Gemini 3.8 Flash 997 276 595 128 4
GPT-6 Astra 950 312 578 156 4
Claude Fable 5.1 862 347 620 169 2
Muse Spark 1.3 <sup>*</sup> 782 306 686 226 0
GLM 5.3 Flash 735 322 731 209 3
GPT-5.6 Sol 673 264 684 378 1
DeepSeek V4.1 Flash 528 328 710 342 92
GPT-5.6 Terra 456 272 718 554 0
GPT-5.6 Luna 430 248 671 606 45

Counts, not cumulative percentages. Each row adds up to 2,000.

Accuracy versus cost #

Muse Spark 1.3 was the cheapest model here and still finished fourth, ahead of five more expensive models. At its Contributor pricing, the entire run cost $0.93 for 782 correct identifications.

Moving to Gemini 3.8 Flash brought that total to 997. That’s 215 more correct answers for another $26.18, an improvement of 10.75 percentage points at roughly 29 times the cost. 💸

Spending beyond that didn’t improve the results. Claude Fable 5.1 cost about twice as much as Gemini, and GPT-6 Astra cost 3.6 times as much. Both scored lower. Gemini delivered the highest accuracy here; Muse offered a cheaper trade-off if identifying fewer species was acceptable.

Estimated totals for 2,000 photos. The dollar axis is logarithmic: spacing represents cost ratios.

<sup>*</sup> Muse Spark 1.3 uses the discounted Contributor tier, which allows provider training/data use.

See the exact numbers #

Accuracy versus estimated cost. Each run contains 2,000 photos.
Model Correct Accuracy Cost
--- --- --- ---
Gemini 3.8 Flash 997 49.85% $27.11
GPT-6 Astra 950 47.50% $97.53
Claude Fable 5.1 862 43.10% $54.19
Muse Spark 1.3 <sup>*</sup> 782 39.10% $0.93
GLM 5.3 Flash 735 36.75% $1.46
GPT-5.6 Sol 673 33.65% $42.89
DeepSeek V4.1 Flash 528 26.40% $8.52
GPT-5.6 Terra 456 22.80% $20.53
GPT-5.6 Luna 430 21.50% $4.99

Is getting half wrong actually bad? #

Getting almost half right is respectable when the task is telling similar species apart. The gap between recognising a family and naming the exact species is where this gets difficult.

Try telling  apart. For *Araniella opisthographa*, [NatureSpot’s identification guidance](https://www.naturespot.org/species/araniella-opisthographa) calls for examination at high magnification. That’s the sort of detail a photo can leave out.

The code and full benchmark results are on [GitHub](https://github.com/qforge-dev/spider-bench).

PS: If you actually know spiders, I’d love to hear what you think on twitter. How obvious are these mistakes to someone who knows what they’re looking at?

P.P.S. Thankfully, I have a cat in charge of spider security at home.

← All postsLabqoat

── more in #computer-vision 4 stories · sorted by recency
── more on @gemini 3.8 flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/i-m-afraid-of-spider…] indexed:0 read:8min 2026-09-21 ·