Ataraxos AI Surpasses Top Stratego Players in Hidden-Information Games

Researchers from MIT, Carnegie Mellon University, New York University, and Stanford University have developed Ataraxos, an AI system that defeated elite human players in Stratego, a board wargame built around hidden information. Ataraxos recorded a 39-2 result against top players at the Stratego world championship and beat the strongest player in the world by 15-1-4.
The work, published in Nature, addresses a difficult class of strategic problems in which participants must act without seeing all relevant information. The researchers say the system achieved higher Stratego performance than prior models while requiring substantially less training computation.
Why Stratego is a demanding AI benchmark
In Stratego, two players arrange 40 pieces and seek to capture the opposing flag. Piece identities remain secret until pieces collide, at which point the lower-ranked piece is removed. That uncertainty makes the value of a move depend on what an opponent may know, believe, or infer rather than on the visible board alone.
The number of possible piece configurations exceeds 10 to the 66th power, making exhaustive consideration impractical. Gabriele Farina, an MIT assistant professor in electrical engineering and computer science and senior author of the paper, said approaches developed for games such as poker could not scale to the number of possible states in Stratego.
Earlier systems, including DeepMind’s DeepNash, used computationally demanding methods and still did not beat the best human Stratego players. Ataraxos was designed to reach superhuman performance without attempting to enumerate every possible future.
Self-play and planning at decision time
The team first trained Ataraxos with self-play reinforcement learning. By repeatedly playing against itself, the model learns a blueprint strategy for arranging pieces and conducting the game. The researchers designed training algorithms intended to learn more efficiently while avoiding the need to predict every possible move.
Farina said Ataraxos reached strictly higher playing strength than DeepNash using less than one hundredth of the training examples and less than one thirtieth of the self-play games. Those figures point to a large efficiency improvement in a setting where training cost had been a major constraint.
During play, the system refines its blueprint through decision-time planning. A generative model estimates the likely identities of an opponent’s concealed pieces, enabling Ataraxos to assess plausible board states and evaluate choices before committing to a move. The researchers describe this planning component as the missing element that enabled superhuman performance.
Potential uses and the need for oversight
The researchers adapted Ataraxos to Barrage Stratego, Hanabi, and Dou dizhu, reporting superhuman performance in each game. The results suggest the method can transfer across imperfect-information games with different rules, including cooperative and competitive formats.
MIT notes that related real-world decisions may arise in business negotiations, cybersecurity, financial markets, and military settings, where parties do not fully know the information or intentions of others. The research team plans to add interpretability measures so people can understand and audit the system’s recommendations.
For businesses evaluating AI in hidden-information workflows, the practical implication is to prioritize systems that can make efficient, auditable recommendations while ensuring that human decision-makers retain final authority over adoption and action.

