{"slug": "ai-pelican-on-a-bike-test-isnt-going-well", "title": "AI “Pelican on a Bike” Test Isn’t Going Well", "summary": "An $80 audit of seven frontier models found that AI labs are gaming the 'Pelican on a Bike' benchmark at the category level, not the individual cell, according to a post by Dylan Castillo. The audit, which used GPT-5.6 Luna to score GPT-5.6 Terra without inter-run reliability checks, highlights the need for cheap external verification of vendor benchmarks.", "body_md": "The first thing that jumped out at me in a post about the famous “[Pelican on a Bike](https://dylancastillo.co/posts/pelicanmaxxing.html)” test is that GPT-5.6 Luna is being used to score GPT-5.6 Terra, without any inter-run reliability check.\n\nIn other words, given a within-lab design, there is a style-level bias test but not a cell-specific bias, which is in fact the thing supposed to be under test.\n\nThe second thing is the entire audit cost $80 across seven frontier models. Independent falsification of a contamination hypothesis is very inexpensive. No cost and all the code and data published means we should be seeing a lot more of this. Cheap external verification is demonstrably feasible, again.\n\nRemember all the noise about Mythos being a marketing scam? Any vendor benchmark that can’t be independently checked is a cynical design decision that deserves heavy pushback and scrutiny.\n\nAnyway, the point of that post seems to be that any lab gaming the Pelican on a Bike benchmark competently games the whole category, not the individual cell. This is the same structure we see in any signature-based detection generally: it catches a weak or clumsy version only.", "url": "https://wpnews.pro/news/ai-pelican-on-a-bike-test-isnt-going-well", "canonical_source": "https://www.flyingpenguin.com/ai-pelican-on-a-bike-test-isnt-going-well/", "published_at": "2026-07-31 07:29:49+00:00", "updated_at": "2026-07-31 07:36:47.286150+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-ethics", "ai-research"], "entities": ["Dylan Castillo", "GPT-5.6 Luna", "GPT-5.6 Terra", "Mythos"], "alternates": {"html": "https://wpnews.pro/news/ai-pelican-on-a-bike-test-isnt-going-well", "markdown": "https://wpnews.pro/news/ai-pelican-on-a-bike-test-isnt-going-well.md", "text": "https://wpnews.pro/news/ai-pelican-on-a-bike-test-isnt-going-well.txt", "jsonld": "https://wpnews.pro/news/ai-pelican-on-a-bike-test-isnt-going-well.jsonld"}}