MIT-Led Team’s AI Ataraxos Achieves Superhuman Stratego Performance A team from MIT, Carnegie Mellon University, New York University, and Stanford University published research in Nature on September 30, 2026 describing Ataraxos, an AI system that achieved superhuman Stratego performance, defeating the strongest human player with a 15-1-4 record and posting a 39-2 record against top players at the world championship. Senior author Gabriele Farina, an assistant professor at MIT's Department of Electrical Engineering and Computer Science, said Ataraxos "reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games." The system combines self-play reinforcement learning with a generative model for decision-time planning, and also reached superhuman performance on Barrage Stratego, Hanabi, and Dou dizhu. September 30, 2026, Inside AI — A collaborative team from MIT , Carnegie Mellon University , New York University , and Stanford University has developed an AI system named Ataraxos that has achieved superhuman performance in the board game Stratego , a complex game of imperfect information. The research, published today in the journal Nature , details how the system defeated top-ranked human players by a significant margin, a feat previously unattainable by AI models. Stratego, often described as military chess, involves hidden piece identities and a massive number of possible game states, exceeding 10 to the 66th power configurations. This complexity has made it a stringent benchmark for AI's strategic reasoning under uncertainty. The new system not only surpassed the previous best AI, DeepMind's DeepNash , but did so with drastically lower computational costs. The breakthrough hinges on a two-pronged approach combining efficient self-play reinforcement learning with a novel decision-time planning mechanism. The researchers trained Ataraxos using self-play to develop a blueprint strategy, then employed a generative model during gameplay to estimate opponent piece identities and refine moves in real time. Gabriele Farina, an assistant professor at MIT's Department of Electrical Engineering and Computer Science and senior author of the paper, explained the challenge: "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting." Read: Anthropic Reveals Hidden Thought Processes in Claude AI The system's efficiency is striking. According to Farina, "Our system reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, indicating a massive improvement in efficiency." This reduction in training resources could make advanced AI more accessible for research and application. Ataraxos demonstrated its prowess by defeating the strongest human Stratego player with a record of 15-1-4 and achieving a 39-2 record against top players at the world championship. Its ability to generalize was also tested on other imperfect information games like Barrage Stratego , Hanabi , and Dou dizhu , where it also achieved superhuman performance. The implications extend beyond board games. Imperfect information scenarios are pervasive in real-world domains such as business negotiations, financial trading, and cybersecurity, where parties operate with incomplete knowledge. The techniques developed for Ataraxos could inform AI systems designed to assist human decision-makers in these areas. Samuel Sokota, a graduate student at Carnegie Mellon and lead author, highlighted the fundamental difference from perfect information games: "The more you bluff, the more your opponent expects it, and the less each bluff is worth. It's not obvious how to reason about that. It's very different from a setting like chess, where the best move is still the best move no matter how often you've played it." The researchers attribute Ataraxos's success partly to its innovative use of a generative model for decision-time planning. Farina noted, "Rather than just guessing blindly, we use decision-time planning to find the most plausible state of the board. Using this generative model allows us to really zoom in on the specific board and opponent we are facing." This capability enables the AI to calculate risk in a composed manner, unlike humans who might overreact when a key piece is threatened. Looking ahead, the team aims to incorporate interpretability measures so that Ataraxos can explain its decisions in human-understandable terms. Farina emphasized the importance of this for adoption: "Humans must have the final say in whether a recommendation is followed, so before adoption can happen, we need a way to audit the model's decisions. We still have a long way to go, but I hope these algorithms can be the foundation for a lot more work to come." Read: AI Model Claude Fable 5 Finds Counterexample to 1939 Math Conjecture The research was supported by the Office of Naval Research , the National Science Foundation , and a Schmidt Sciences AI2050 Early Career Fellowship , among others. As AI continues to advance in handling uncertainty, systems like Ataraxos may pave the way for more robust and efficient decision-support tools in complex, real-world scenarios.