AI Research4 mins read

Ataraxos AI beats Stratego’s greatest player in major board-game milestone

Researchers from Carnegie Mellon, NYU, Stanford, and MIT built Ataraxos, an AI system that decisively beat top Stratego player Pim Niemeijer, highlighting progress in imperfect-information games.

The result: Ataraxos defeats a Stratego legend

Ataraxos, an AI system from researchers at Carnegie Mellon University, New York University, Stanford University, and MIT, beat Dutch player Pim Niemeijer in an official 20-game Stratego series. The final score was 15 wins, one loss, and four draws for the AI.

Niemeijer is described in the article as the most decorated player in Stratego history, with multiple world, national, and online titles. The outcome marks a major step for AI in a game where hidden information has long protected human advantage.

Why Stratego is harder than it looks

Stratego setup and hidden-piece gameplay illustration
Image credits:Sokota, S., Vinitsky, E., Hu, H. et al.

Stratego is difficult for AI because both players set up 40 pieces face down, and a piece’s rank is revealed only through contact with an opponent’s piece. The article says there are more than 10^33 possible setups, making the game a demanding test of decisions under uncertainty.

Unlike perfect-information games, a move’s value depends on what may have happened before and what each player may know or infer. That makes Stratego a strong benchmark for reinforcement learning and search methods designed for imperfect-information environments.

A low-cost academic system outperformed prior efforts

The article contrasts Ataraxos with DeepMind’s earlier DeepNash system, which did not surpass the level of top human Stratego players. According to the article, DeepNash’s training would be estimated at $3 million to $4.5 million at 2025 prices.

Ataraxos reportedly trained for less than $8,000, using one week on 16 Nvidia H100 GPUs plus four more days on four GPUs for a belief network. The researchers credit a custom GPU simulator and higher sample efficiency for the lower compute cost.

How Ataraxos learned to handle hidden information

Ataraxos trained without human game data, learning through self-play. A key technique described in the article is regularization, which pushes the AI to vary its setups and moves instead of settling too early into predictable patterns.

The system also uses a belief network to predict hidden opponent pieces, generate possible game states, and evaluate candidate moves before acting. This combination helped the AI stay unpredictable in a game where bluffing, uncertainty, and inference matter.

Why the milestone matters beyond board games

The same method was also applied to Barrage Stratego, Hanabi, and Dou dizhu, according to the article, with strong results across those games. The researchers argue that large amounts of hidden information are no longer necessarily a barrier for reinforcement learning and search.

The broader takeaway is practical: AI systems may become more useful in strategic decision problems where participants have incomplete information, provided fast and accurate simulators can be built. The researchers also note a limitation: the current search approach does not keep improving simply by adding more compute time.

Discover More