Is the model actually getting dumber, or are we just reading tea leaves from single samples? A developer involved in building Folkbench argues that single-sample judgments of LLM degradation, such as the viral pelican SVG test, are statistically meaningless because LLMs are probabilistic samplers with randomness baked in. The developer recommends fixing model version, reasoning effort, prompt and time window, then comparing pass rates across many runs — for example 16 of 20 versus 8 of 20 — before concluding a model has been watered down. Disclosure: I'm involved in building Folkbench, which I mention near the end of this post. I've been chewing on a question that's been discussed to death but never really settled: when we say a model "got dumber" or "got nerfed," how much of that is real, and how much is just the illusion of a single sample? Lately a lot of people I know have been throwing the pelican test at models have it draw an SVG of a pelican riding a bicycle . It's a genuinely brutal test — it's not just about writing code, it's about spatial reasoning: beak and handlebars, feet and pedals, frame and wheels. If the model's spatial understanding is even slightly off, you get something pretty abstract. But after running it a few dozen times, I think the biggest trap with the pelican test is judging by a single output. An LLM is a probabilistic sampler with randomness baked in. Same prompt: on one run the spatial awareness is maxed out — frame, cranks, foot placement all correct; run it again later and suddenly it's postmodern abstract art. If provider A gets a basically sensible pose in 16 out of 20 runs and provider B only passes 8 out of 20, then the difference means something statistically. Declaring "A is the full model, B is watered down" based on one random screenshot is basically flipping a coin. To actually compare anything, you need a shared baseline — fix the model version, reasoning effort, prompt and time window, set a reference point first, and then look only at relative differences. That leads to a few really painful engineering details: I'd been running these tests by hand for a while, and the biggest pain was the cost of record-keeping: repeated runs, saving SVGs, logging parameters, aligning timestamps — after a few dozen batches you're numb. What eats the time usually isn't the test itself but all the tedious logging and organizing around it. That said, if you just want a quick read on how different models perform, you don't have to do it all yourself — there's a ready-made leaderboard on Folkbench https://folkbench.com/?utm source=luntan&utm campaign=dev https://folkbench.com/?utm source=luntan&utm campaign=dev that can save you the effort. If you're also playing with this test, I'd love to hear how you run it — or see your most cursed results.