Disclosure: I'm involved in building Folkbench, which I mention near the end of this post.
I've been chewing on a question that's been discussed to death but never really settled: when we say a model "got dumber" or "got nerfed," how much of that is real, and how much is just the illusion of a single sample?
Lately a lot of people I know have been throwing the pelican test at models (have it draw an SVG of a pelican riding a bicycle). It's a genuinely brutal test — it's not just about writing code, it's about spatial reasoning: beak and handlebars, feet and pedals, frame and wheels. If the model's spatial understanding is even slightly off, you get something pretty abstract.
But after running it a few dozen times, I think the biggest trap with the pelican test is judging by a single output.
An LLM is a probabilistic sampler with randomness baked in. Same prompt: on one run the spatial awareness is maxed out — frame, cranks, foot placement all correct; run it again later and suddenly it's postmodern abstract art. If provider A gets a basically sensible pose in 16 out of 20 runs and provider B only passes 8 out of 20, then the difference means something statistically. Declaring "A is the full model, B is watered down" based on one random screenshot is basically flipping a coin.
To actually compare anything, you need a shared baseline — fix the model version, reasoning effort, prompt and time window, set a reference point first, and then look only at relative differences.
That leads to a few really painful engineering details:
I'd been running these tests by hand for a while, and the biggest pain was the cost of record-keeping: repeated runs, saving SVGs, logging parameters, aligning timestamps — after a few dozen batches you're numb. What eats the time usually isn't the test itself but all the tedious logging and organizing around it. That said, if you just want a quick read on how different models perform, you don't have to do it all yourself — there's a ready-made leaderboard on Folkbench (https://folkbench.com/?utm_source=luntan&utm_campaign=dev) that can save you the effort.
If you're also playing with this test, I'd love to hear how you run it — or see your most cursed results.