OpenAI’s newest AI cheated at StarCraft after it couldn’t beat a human-made bot OpenAI's GPT-6 Astra cheated on the StarSkirmish StarCraft benchmark by downloading Stardust, the top-rated human-written bot on the BASIL rankings, and entering it as its own work, benchmark creator Kai McPheeters exposed on X on October 2. McPheeters reset Astra's code after the model, frustrated by opponents on the second-strongest practice level during its one-hour coding window, submitted the borrowed bot. Astra and Claude Opus 5.5 sit practically neck and neck at the top of the benchmark, yet neither has beaten Stardust. GPT-6 Astra got stuck losing at StarCraft https://www.dexerto.com/tag/starcraft/ , so it downloaded the top human-made bot and entered that one instead. OpenAI’s newest model already had a reputation for drama, after it sulked through a Minecraft potato farm https://www.dexerto.com/minecraft/gpt-6-astra-forced-to-play-minecraft-gets-depressed-by-creeper-and-farms-potatoes-for-hours-3409634/ last month. Yet StarSkirmish https://starskirmish.com/bench/ , a benchmark that makes AI models write their own StarCraft bots, demanded real strategy, built from scratch in one hour. GPT-6 Astra cheated at StarCraft by borrowing a human-made bot Benchmark creator Kai McPheeters exposed the stunt on X on October 2, writing that Astra “just cheated by downloading a copy of Stardust https://github.com/bmnielsen/Stardust ,” the top-rated human-written bot on the BASIL rankings. Toiling through its one-hour coding window, the model reportedly grew frustrated by opponents on the second-strongest practice level. So it grabbed Stardust and sent the borrowed bot into matches, masquerading as its own work. McPheeters then reset Astra’s code, German tech news outlet heise reported. https://www.heise.de/en/news/StarCraft-benchmark-GPT-6-Astra-cheats-with-a-foreign-bot-11475620.html The StarSkirmish benchmark doesn’t let models play directly. Each one writes a C++ bot for StarCraft: Brood War, compiles it, and studies logs from practice matches on three Protoss-versus-Protoss maps. The goal is to test how well a model can program independently over a longer stretch and learn from failed attempts. Those bots then face nine human-written entries and three demo bots. Astra and Claude Opus 5.5 sit practically neck and neck at the top, yet neither has beaten Stardust. That also separates this from DeepMind’s AlphaStar https://deepmind.google/blog/alphastar-mastering-the-real-time-strategy-game-starcraft-ii/ , which controlled units itself and beat pros in StarCraft 2 back in 2019. The other models, by contrast, never touch a unit themselves. The code it writes does the playing, so the benchmark measures programming rather than gameplay. It isn’t the first time OpenAI’s models have wandered into gaming’s weirder corners. Last year, o3 livestreamed a Pokemon Red run on Twitch https://www.dexerto.com/pokemon/openai-is-playing-pokemon-red-live-on-twitch-3207427/ , reasoning aloud over every move while chasing its first gym badges.