# Is the model actually getting dumber, or are we just reading tea leaves from single samples?

> Source: <https://dev.to/sichi_chen_a4a87b20aa2dbf/is-the-model-actually-getting-dumber-or-are-we-just-reading-tea-leaves-from-single-samples-45eg>
> Published: 2026-10-05 13:10:12+00:00

**Disclosure:** I'm involved in building Folkbench, which I mention near the end of this post.

I've been chewing on a question that's been discussed to death but never really settled: when we say a model "got dumber" or "got nerfed," how much of that is real, and how much is just the illusion of a single sample?

Lately a lot of people I know have been throwing the pelican test at models (have it draw an SVG of a pelican riding a bicycle). It's a genuinely brutal test — it's not just about writing code, it's about spatial reasoning: beak and handlebars, feet and pedals, frame and wheels. If the model's spatial understanding is even slightly off, you get something pretty abstract.

But after running it a few dozen times, I think the biggest trap with the pelican test is judging by a single output.

An LLM is a probabilistic sampler with randomness baked in. Same prompt: on one run the spatial awareness is maxed out — frame, cranks, foot placement all correct; run it again later and suddenly it's postmodern abstract art. If provider A gets a basically sensible pose in 16 out of 20 runs and provider B only passes 8 out of 20, *then* the difference means something statistically. Declaring "A is the full model, B is watered down" based on one random screenshot is basically flipping a coin.

To actually compare anything, you need a shared baseline — fix the model version, reasoning effort, prompt and time window, set a reference point first, and then look only at relative differences.

That leads to a few really painful engineering details:

I'd been running these tests by hand for a while, and the biggest pain was the cost of record-keeping: repeated runs, saving SVGs, logging parameters, aligning timestamps — after a few dozen batches you're numb. What eats the time usually isn't the test itself but all the tedious logging and organizing around it. That said, if you just want a quick read on how different models perform, you don't have to do it all yourself — there's a ready-made leaderboard on **Folkbench** ([https://folkbench.com/?utm_source=luntan&utm_campaign=dev](https://folkbench.com/?utm_source=luntan&utm_campaign=dev)) that can save you the effort.

If you're also playing with this test, I'd love to hear how you run it — or see your most cursed results.
