# AI “Pelican on a Bike” Test Isn’t Going Well

> Source: <https://www.flyingpenguin.com/ai-pelican-on-a-bike-test-isnt-going-well/>
> Published: 2026-07-31 07:29:49+00:00

The first thing that jumped out at me in a post about the famous “[Pelican on a Bike](https://dylancastillo.co/posts/pelicanmaxxing.html)” test is that GPT-5.6 Luna is being used to score GPT-5.6 Terra, without any inter-run reliability check.

In other words, given a within-lab design, there is a style-level bias test but not a cell-specific bias, which is in fact the thing supposed to be under test.

The second thing is the entire audit cost $80 across seven frontier models. Independent falsification of a contamination hypothesis is very inexpensive. No cost and all the code and data published means we should be seeing a lot more of this. Cheap external verification is demonstrably feasible, again.

Remember all the noise about Mythos being a marketing scam? Any vendor benchmark that can’t be independently checked is a cynical design decision that deserves heavy pushback and scrutiny.

Anyway, the point of that post seems to be that any lab gaming the Pelican on a Bike benchmark competently games the whole category, not the individual cell. This is the same structure we see in any signature-based detection generally: it catches a weak or clumsy version only.
