Partner ContentInside Arkadium’s GameLab, PhDs from MIT and Harvard are breaking the world’s most advanced AI models with games Arkadium's GameLab, featuring researchers from MIT and Harvard, is using games to expose weaknesses in AI models, with tests showing ChatGPT making the same invalid move 16 times in Block Champ. Josh Tenenbaum, professor at MIT, and Lance Ying, with PhDs from Harvard and MIT, argue that games provide a more complete benchmark for machine intelligence than traditional tests. Kenny Rosenblatt, CEO of Arkadium, highlighted the opportunity to apply the company's game catalog to understanding AI reasoning and problem-solving. Become a member of GB MAX to gain exclusive access to the industry and to the most influential global B2B leadership community in the business of gaming, entertainment, and tech. Join now https://go.gamesbeat.com/gb-max/ and also get a VIP ticket to GamesBeat Next Nov 2-3, SF . Presented by Arkadium For decades, games have challenged humans to think strategically, solve problems and adapt to new situations. Now, researchers at MIT and Harvard working with Arkadium https://gamesbeat.com/arkadium-launches-gamelab-to-make-ai-ready-for-the-real-world/ believe those same games could become one of artificial intelligence’s most valuable testing grounds. Through Arkadium’s GameLab initiative https://gamesbeat.com/arkadium-launches-gamelab-to-make-ai-ready-for-the-real-world/ , researchers including Josh Tenenbaum, professor of computational cognitive science at MIT, and Lance Ying, who holds PhDs from Harvard and MIT, are exploring how games can serve as benchmarks for reasoning, planning, memory and decision-making. Rather than relying on traditional AI benchmarks, they believe the diversity of human games offers a more complete picture of machine intelligence. “We believe that the multiverse of human games is an ideal testbed for measuring machine intelligence,” said Tenenbaum. “Games are powerful cultural artifacts, designed to be effective miniatures and abstractions of real-world human activities, problems, challenges, enterprises, and dynamics for training and preparing humans for adaptation and problem solving in the real world. Games cover nearly every human skill, from strategic planning and resource management to social interaction and deception, pattern recognition, and navigating complex physical environments.” While today’s AI models can generate code, solve complex math problems, answer questions and produce human-like text, they continue to struggle with many of the reasoning challenges humans encounter in everyday games. Current systems often excel at recognizing patterns and predicting likely outcomes but struggle when they need to maintain context, adapt to changing information or make decisions without a clear answer. Through GameLab, Arkadium is using its catalog of games to explore those gaps, testing AI capabilities including spatial reasoning, long-term planning, memory, constraint management and decision-making under uncertainty. “Our games have always been designed to challenge people to think, adapt and solve problems,” said Kenny Rosenblatt, CEO of Arkadium. “Now, we have the opportunity to apply that same foundation to understanding AI, using the breadth of our catalog to help researchers explore how machines reason, learn and navigate complex challenges.” Block Champ reveals AI’s spatial reasoning challenge Arkadium’s GameLab testing shows how simple games can expose complex reasoning gaps. In Block Champ, players place puzzle pieces on a board while planning several moves ahead. Success requires understanding spatial relationships and adapting as the board changes. During testing, ChatGPT repeatedly attempted to place a block in a space where it could not fit, making the same invalid move 16 times in a row despite receiving feedback. AI models often struggle to maintain an accurate representation of the environment they are interacting with. “The model often fails to maintain an accurate representation of the grid and to check whether a new piece actually fits,” said Ying. “Doing that reliably requires simulating putting pieces onto the game grid and verifying it against the constraints.” The example highlights a challenge in AI development: describing a problem and reasoning through a problem are not always the same thing. Daily Crossword exposes AI’s constraint problem Crossword puzzles test more than vocabulary. Players must combine knowledge with strict rules, continuously updating their assumptions as new answers provide additional information. In one GameLab test, Claude Opus 4.7 repeatedly attempted to place a 12-letter answer into a 10-letter space. Even after identifying the mistake, the model continued making similar errors. The challenge was applying the puzzle’s constraints throughout the solving process. “Coming up with candidate answers from clues and figuring out the letter constraint from the crossword spatial inputs are two different cognitive processes,” said Ying. “Retrieving a word that matches the description is easy associative recall,” he continued. “Recognizing and enforcing the spatial constraint in problem solving is a separate cognitive process that remains challenging for today’s AI models.” Gin Rummy tests strategic decision-making Card games introduce another layer of complexity: making decisions when information is incomplete. In Gin Rummy, players must weigh risk and reward, anticipate an opponent’s possible actions and adjust their strategy without knowing the full state of the game. That uncertainty is what makes these environments valuable for AI research. The challenge is evaluating possibilities, managing uncertainty and knowing when information is incomplete. “Today’s AI models struggle with Gin Rummy because they’re often rewarded for being confident,” said Ying. “As a result, they make strong but incorrect assumptions and hallucinate information in their decision making where they should stay uncertain.” Why games matter for AI research For Arkadium, the goal is to better understand the gap between today’s AI systems and the kinds of reasoning required in the real world. Games provide structured environments where planning, memory, adaptation and decision-making can be measured at scale. Unlike traditional cognitive experiments, which often capture a single moment of human behavior, games reveal how strategies develop and evolve over time. “The longitudinal aspect of learning and problem solving is largely ignored in traditional cognitive experiments,” said Ying. “Many games are longitudinal. They allow us to observe and experiment with how complex strategies form, solidify and evolve over months or even years.” Researchers believe large-scale gameplay data could eventually help reveal more about intelligence itself. “Analyzing large-scale gameplay data helps to inform the cognitive taxonomy and architecture of human intelligence,” said Tenenbaum. “Because internet-scale gaming offers incredibly diverse, multi-faceted environments with large-scale human problem-solving trajectories, analyzing this behavioral data could allow us to probe the underlying structure of human intelligence.” Tenenbaum believes one of the simplest ways to measure AI progress may be through the games millions of people already play every day. “One intuitive test of AI progress would be to look at the top charts on the iOS App Store or Steam and ask: how many of these games can today’s models actually learn and play?” he said. “These charts span an enormous variety of genres, interfaces, rules and social settings, and they change constantly as new games are released and existing games are updated. Rather than optimizing for a fixed collection of benchmark questions, AI systems would have to continually adapt to unfamiliar environments and challenges.” As AI moves into higher-stakes industries, understanding these limitations will become vital. If models still struggle with a crossword puzzle, a spatial challenge or a card game, those failures reveal where additional progress is needed before AI can reliably handle more complex real-world decisions.