
LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.
Researchers from Carnegie Mellon, NYU, Stanford, and MIT built Ataraxos, an AI system that decisively beat top Stratego player Pim Niemeijer, highlighting progress in imperfect-information games.

Ataraxos, an AI system from researchers at Carnegie Mellon University, New York University, Stanford University, and MIT, beat Dutch player Pim Niemeijer in an official 20-game Stratego series. The final score was 15 wins, one loss, and four draws for the AI.
Niemeijer is described in the article as the most decorated player in Stratego history, with multiple world, national, and online titles. The outcome marks a major step for AI in a game where hidden information has long protected human advantage.

Stratego is difficult for AI because both players set up 40 pieces face down, and a piece’s rank is revealed only through contact with an opponent’s piece. The article says there are more than 10^33 possible setups, making the game a demanding test of decisions under uncertainty.
Unlike perfect-information games, a move’s value depends on what may have happened before and what each player may know or infer. That makes Stratego a strong benchmark for reinforcement learning and search methods designed for imperfect-information environments.
The article contrasts Ataraxos with DeepMind’s earlier DeepNash system, which did not surpass the level of top human Stratego players. According to the article, DeepNash’s training would be estimated at $3 million to $4.5 million at 2025 prices.
Ataraxos reportedly trained for less than $8,000, using one week on 16 Nvidia H100 GPUs plus four more days on four GPUs for a belief network. The researchers credit a custom GPU simulator and higher sample efficiency for the lower compute cost.
Ataraxos trained without human game data, learning through self-play. A key technique described in the article is regularization, which pushes the AI to vary its setups and moves instead of settling too early into predictable patterns.
The system also uses a belief network to predict hidden opponent pieces, generate possible game states, and evaluate candidate moves before acting. This combination helped the AI stay unpredictable in a game where bluffing, uncertainty, and inference matter.
The same method was also applied to Barrage Stratego, Hanabi, and Dou dizhu, according to the article, with strong results across those games. The researchers argue that large amounts of hidden information are no longer necessarily a barrier for reinforcement learning and search.
The broader takeaway is practical: AI systems may become more useful in strategic decision problems where participants have incomplete information, provided fast and accurate simulators can be built. The researchers also note a limitation: the current search approach does not keep improving simply by adding more compute time.

LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.

RRSI aims to help self-improving AI agents generalize beyond their test tasks while using fewer tokens.

An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

New task-log analysis shows AI agents expand what researchers can attempt, but autonomy remains limited.