A Stratego AI beat the world's best 15-1 on 16 GPUs, not 1,024 chips
Ataraxos beat Pim Niemeijer 15 games to one, and its training cost is the bigger story. What the result proves, and what it doesn't.
Researchers from Carnegie Mellon, MIT, NYU and Stanford say their AI, Ataraxos, beat Pim Niemeijer, arguably the best Stratego player in history, 15 games to one with four draws. According to the team, it was trained on 16 GPUs for a week, plus four more GPUs for four days. DeepMind's earlier attempt, DeepNash, used 1,024 specialized chips for two to three months, a run the Ataraxos team estimates at $3 million to $4.5 million at 2025 prices.
Why Stratego Resisted
Stratego is an imperfect-information game: you can see where the opponent's 40 pieces are, but not what they are. Identities are revealed only when pieces collide. Poker is also imperfect-information, but Texas Hold'em has just 1,326 possible hands, few enough to weigh exhaustively. Stratego has more than a decillion possible setups, and a game can run 2,000 moves against roughly 40 in chess.
Bluffing makes it worse. Move a weak piece as if it were a marshal and you may scare the opponent off, but bluff too often and your threats mean nothing; never bluff and you are predictable. The researchers say that balance is what defeated earlier systems.
How Ataraxos Works
Like DeepNash, Ataraxos learned by self-play, 163 million games in total: moves that led to wins were played more often, moves that led to losses less. Two changes mattered. First, hidden information tends to send self-play learning in circles, so the team made large strategy adjustments early in training and small ones later.
Second, it can think ahead before each move, which DeepNash never did. Search of this kind, used by AlphaGo, was considered too expensive in Stratego because there were too many possible boards to examine. The fix was a second neural network, a belief model, that guesses the opponent's hidden pieces from how they have moved. Instead of enumerating every arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on results.
What Is Demonstrated, and What Isn't
Demonstrated: a peer-reviewed result (Nature, 2026) against a four-time world champion, plus 38 wins in 40 games against attendees at the 2025 World Championship. The same approach also beat three world champions at Barrage Stratego, mastered the cooperative card game Hanabi, and beat the best bots at dou dizhu.
Caveats the source itself supplies: the Niemeijer match was 20 online games over three weeks, a small sample in a game where luck is built in. The researchers say even a perfect strategy sometimes loses. Niemeijer also knew the AI would not adapt to him, which cuts both ways: it gave him time to probe for weaknesses, but it also means the test never covered an opponent the AI had to adjust to.
Not demonstrated: anything outside games. The team suggests war gaming, negotiation and financial markets as future targets, and argues the gap is smaller than it looks because any real problem starts with a simplified model. That is an argument, not a result. Real problems lack fixed rules and clear winners, and nobody has shown Ataraxos working on one. The system also can't explain its moves; Farina said the team is not yet at strong, interpretable strategies.
Questions You Should Be Asking
- If the efficiency gain came largely from a custom simulator running millions of moves per second on GPUs, how much of it transfers to a domain where no such simulator exists?
- The belief model guesses hidden information from observed behavior. What happens when the opponent's behavior looks nothing like the self-play opponents it trained against?
- Twenty games is a small sample. How many would it take to separate a real edge from luck, and would the 15-1 margin survive an opponent who could adapt?
- For real-world uses like negotiation or conflict simulation, who builds the simplified model, and who checks that it reflects reality rather than the builder's assumptions?
- If the system cannot explain its moves, how would you audit a recommendation before acting on it?
What To Watch Next
The signal is whether the approach leaves games. Watch for a published application where the rules are not fixed and no cheap simulator exists, and for progress on the explainability gap Farina acknowledged. Until then, the solid result is cheaper, stronger play in hidden-information games, not a general tool for negotiation or markets.
- 1Optimize AI training efficiency by using specialized algorithms instead of raw computing power to reduce costs by 99%+.
- 2Leverage imperfect-information game strategies to improve decision-making systems when data is incomplete or hidden.
- 3Benchmark AI performance against human experts in complex domains to validate real-world applicability beyond synthetic tests.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
