AI “Pelican on a Bike” Test Isn’t Going Well An $80 audit of seven frontier models found that AI labs are gaming the 'Pelican on a Bike' benchmark at the category level, not the individual cell, according to a post by Dylan Castillo. The audit, which used GPT-5.6 Luna to score GPT-5.6 Terra without inter-run reliability checks, highlights the need for cheap external verification of vendor benchmarks. The first thing that jumped out at me in a post about the famous “ Pelican on a Bike https://dylancastillo.co/posts/pelicanmaxxing.html ” test is that GPT-5.6 Luna is being used to score GPT-5.6 Terra, without any inter-run reliability check. In other words, given a within-lab design, there is a style-level bias test but not a cell-specific bias, which is in fact the thing supposed to be under test. The second thing is the entire audit cost $80 across seven frontier models. Independent falsification of a contamination hypothesis is very inexpensive. No cost and all the code and data published means we should be seeing a lot more of this. Cheap external verification is demonstrably feasible, again. Remember all the noise about Mythos being a marketing scam? Any vendor benchmark that can’t be independently checked is a cynical design decision that deserves heavy pushback and scrutiny. Anyway, the point of that post seems to be that any lab gaming the Pelican on a Bike benchmark competently games the whole category, not the individual cell. This is the same structure we see in any signature-based detection generally: it catches a weak or clumsy version only.