We made Grok 4.5, GPT-5.5, and Claude build the same apps
XAI's Grok 4.5, OpenAI's GPT-5.5, and Anthropic's Claude Opus 4.8 and Fable 5 were benchmarked on one-shot app generation across three interactive tasks. Claude models won the Rubik's Cube round, GPT-5.5 won the gravity …