Could World of Goo tower building become a useful benchmark for AI agents? A challenge called Goo AI Arena has opened, asking participants to use AI to build the tallest standing tower in World of Goo using all 300 balls through legal in-game building steps. The first entry, built by the author with Codex, reached 13.9 metres and was titled "It's a tower," with the author noting that teaching AI is hard and that separating model capability from tooling, compute budget, and human assistance remains the harder question for using the task as an agent benchmark. I’ve opened Goo AI Arena , a challenge to use AI to build the tallest standing tower in World of Goo. The task: use all 300 balls through legal in-game building steps and keep the finished structure standing. There’s a leaderboard, submission instructions, and a completed example. Codex and I built the first entry: 13.9 metres , titled “It’s a tower.” I provided guidance, so this is an example of human–AI collaboration. My main lesson: teaching AI is hard. Beyond the competition, I’m interested in its potential for agent evaluation. Building a tower involves planning, interpreting a changing physical environment, and recovering when construction doesn’t go as expected. Height is an easy outcome to measure; separating model capability from tooling, compute budget, and human assistance is the harder question. I’d welcome both new contenders and ideas for making comparisons scientifically useful . What controls would you want before treating this as a benchmark?