An artificial intelligence system developed by researchers at MIT, Carnegie Mellon University, New York University, and Stanford University has accomplished something no machine had managed before: defeating the world’s top-ranked human players at the board wargame Stratego, and doing so by a record margin. The system, called Ataraxos, achieved a 15-1-4 score against the strongest Stratego player on the planet and posted a 39-2 record against elite human competitors at the Stratego world championship. Remarkably, it reached this level of play while consuming a tiny fraction of the computational resources required by earlier attempts, outperforming models that cost millions of dollars to train. The research, published in Nature, signals a major advance in the long-standing challenge of teaching machines to reason strategically when crucial information is hidden from view.
Stratego has long served as a benchmark for testing the strategic thinking abilities of powerful AI models, and for good reason. The two-player game resembles military chess: each player arranges 40 pieces on their side of the board and moves them across it with the goal of capturing the opponent’s flag. But unlike chess, the identity of every piece remains secret until two pieces collide, at which point the lower-ranking piece is eliminated. This structure makes Stratego a game of imperfect information, a category of problems that has historically proven far more difficult for machines than perfect-information games. The number of possible piece configurations exceeds 10 to the 66th power, an exponentially greater figure than in chess, creating a search space so vast that brute-force approaches collapse under their own weight.
The difficulty is not merely computational but conceptual. In games with hidden information, the value of any given action depends on what an opponent believes and expects, creating a web of intertwined decisions that resists straightforward analysis. Samuel Sokota, a graduate student at Carnegie Mellon and lead author of the paper, illustrated the problem with the example of bluffing. The more a player bluffs, the more the opponent comes to expect it, and the less each individual bluff is worth. This dynamic is fundamentally different from chess, where the best move remains the best move no matter how often it has been played. In imperfect information settings, there is no fixed table of optimal actions; strategy must adapt continuously to an evolving psychological landscape.
Earlier efforts to crack Stratego illustrated just how steep the challenge was. Google DeepMind developed a system called DeepNash that relied on sophisticated but computationally demanding and costly operations. Even with training costs running into millions of dollars, those models were still not strong enough to beat top human Stratego players. Gabriele Farina, an assistant professor in MIT’s Department of Electrical Engineering and Computer Science and a principal investigator at the Laboratory for Information and Decision Systems, who is senior author of the new paper, explained the core obstacle. With Stratego, he noted, there is an explosion of possible universes a player might have to deal with, and AI techniques developed for games like poker simply could not scale to this setting.
The MIT-led team set out to build a full AI system capable of superhuman performance at dramatically lower cost, and they named it Ataraxos, after a Greek word describing a person who is unbothered or free from anxiety. The name proved apt. The researchers built the system on a foundation of self-play reinforcement learning, a training technique in which the model plays against itself many times to gradually learn a strong blueprint strategy for excelling at the game. What distinguished their approach was the design of especially efficient algorithms that allowed Ataraxos to learn far faster than prior methods while avoiding the trap of trying to predict every possible move, a pitfall that had crippled earlier systems in Stratego’s enormous state space.
The efficiency gains were striking. According to Farina, Ataraxos reaches strictly higher playing strength than DeepNash while using less than one hundredth of the training examples and less than one thirtieth of the self-play games, which he characterized as a massive improvement in efficiency. This means the system not only plays better but does so with orders of magnitude less data and computation, a combination that has rarely been achieved in competitive AI benchmarks. The result suggests that clever algorithmic design, rather than sheer scale of resources, may be the key to unlocking superhuman performance in the hardest classes of strategic games.
Training, however, was only half of the two-pronged approach. During an actual game, Ataraxos uses its learned blueprint strategy as a starting point to set up the board and begin considering its next moves at each round of play. Before acting, though, it refines its choices on the fly using a technique called decision-time planning. The system employs a generative model that uses probabilities to estimate the likely identities of the opponent’s hidden pieces, then evaluates future choices before committing to a move. Rather than guessing blindly, Farina explained, the team uses decision-time planning to find the most plausible state of the board, and the generative model allows the system to zoom in on the specific board and opponent it is facing.
This innovative use of a generative model for decision-time planning was, in the researchers’ view, the missing piece that finally enabled superhuman performance. The system’s composure under pressure became a defining feature of its play. Farina observed that Ataraxos is good at calculating risk in a way that humans are not. A human player might start to panic if their most valuable piece is exposed, but the bot can be surprisingly composed, refusing to overcorrect and give away its secrets. That emotional steadiness, backed by probabilistic reasoning about what the opponent likely holds, translated directly into the lopsided scores the system recorded against the world’s best players.
Crucially, the researchers demonstrated that Ataraxos is not a one-game specialist. They adapted the system for other imperfect information games with very different rules and designs, including Barrage Stratego, a faster-paced variant with fewer pieces; Hanabi, a cooperative card game with many players; and Dou dizhu, a game in which two players cooperate against a third. In each instance the system achieved superhuman performance, evidence that the underlying method generalizes across a variety of use cases rather than exploiting quirks of a single ruleset. That generality matters enormously for the technology’s prospects beyond the game board, because real-world strategic problems rarely come in one standardized format.
The broader ambition is to carry these capabilities into situations where humans must make strategic decisions with incomplete knowledge, such as business negotiations, cybersecurity, or military maneuvers. The world is full of imperfect information problems: traders in financial markets may not know the rationale behind the trades of others, and military forces rarely have full knowledge of enemy positions. In such interactions, the decisions parties make, and the decisions they choose not to make, are intertwined in ways that make it extremely difficult to determine the best next step. Farina emphasized that in the imperfect information tasks one faces in reality, there is often no luxury of enumerating all the possibilities because there are simply too many, and having general-purpose AI algorithms that can provably perform this challenging task so well represents a big step forward.
Significant work remains before such systems can be trusted with consequential recommendations. The researchers’ next goal is to build interpretability measures into Ataraxos so the system can explain its decision-making in a way a human could understand. Farina stressed that humans must have the final say in whether a recommendation is followed, and that before adoption can happen, there needs to be a way to audit the model’s decisions. He acknowledged there is still a long way to go, but expressed hope that these algorithms can serve as the foundation for much more work to come. The research was funded in part by the Office of Naval Research, the New York University Department of Civil and Urban Engineering, the C2SMART Center, the National Science Foundation, and a Schmidt Sciences AI2050 Early Career Fellowship. Alongside Farina and Sokota, the author team includes Eugene Vinitsky and Zico Kolter of New York University, Hengyuan Hu of Stanford, and Zhiyuan Fan, an EECS graduate student at MIT.
Subject of Research: An AI system achieving superhuman performance in Stratego and other hidden-information games through efficient self-play reinforcement learning and decision-time planning.
Article Title: This game-playing AI is the new champ at Stratego
Article References: This game-playing AI is the new champ at Stratego. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: artificial intelligence, Stratego, imperfect information games, reinforcement learning, self-play, decision-time planning, generative models, MIT, Nature, game theory, strategic decision making, DeepNash
Cite Scienmag News
Courtney Benton. (September 30, 2026). New AI conquers Stratego, beating world champions at hidden-information warfare. Scienmag. https://scienmag.com/new-ai-conquers-stratego-beating-world-champions-at-hidden-information-warfare/
Courtney Benton. "New AI conquers Stratego, beating world champions at hidden-information warfare." Scienmag, 30 September 2026, https://scienmag.com/new-ai-conquers-stratego-beating-world-champions-at-hidden-information-warfare/. Accessed 30 September 2026.
Courtney Benton. "New AI conquers Stratego, beating world champions at hidden-information warfare." Scienmag. September 30, 2026. https://scienmag.com/new-ai-conquers-stratego-beating-world-champions-at-hidden-information-warfare/

