On September 30, researchers from MIT, Carnegie Mellon University (CMU), New York University and Stanford published a paper in Nature describing an AI system called Ataraxos. In the strategy board game Stratego, it beat Pim Niemeijer — widely regarded as the most accomplished human player in the game's history — with a record of 15 wins, 1 loss and 4 draws, a margin that left little doubt about which side had the upper hand across the match.
The cost was strikingly low. Per the configuration described in the paper, the reinforcement learning model was trained on 16 H100 GPUs for one week, and a belief model used to infer the identity of an opponent's hidden pieces was trained separately on 4 H100s for four days. At 2025 prices, the total bill came to under $8,000 — roughly what a university lab might spend on a single grad-student conference trip, not a dedicated game-AI research effort.
Why Stratego is hard
Stratego is a two-player wargame in which each side commands 40 pieces, ranked from a lowly scout up to a marshal, alongside bombs and a flag that has to be defended. At the start, every piece faces away from the opponent — who holds the marshal, which pieces are bombs, and where the flag is hidden are all invisible until a piece is directly challenged in combat. The research team estimates the number of possible piece arrangements exceeds 10 to the 66th power, a search space far too large to brute-force.
That sets Stratego apart from Go and chess, which are games of perfect information — everything on the board is visible to both players at all times, so the challenge is purely about calculating ahead. Stratego requires guessing at an opponent's hidden setup while also managing how much of your own information you give away through your moves, and knowing when to bluff rather than play it safe. Games of imperfect information like this are structurally closer to real-world problems such as negotiation, auctions and cyber conflict, where each side acts on incomplete knowledge of the other, and they have long been a hard problem for AI research precisely because there is no single correct board state to plan against.
Gabriele Farina, the paper's senior author and a professor in MIT's Department of Electrical Engineering and Computer Science, described Ataraxos's style of play this way:
"Ataraxos is good at calculating risk in a way that humans are not... doesn't overcorrect and give away its secrets."
Measured against DeepNash
The previous benchmark in Stratego was DeepMind's DeepNash, released in 2022. Trained through large-scale self-play, it climbed into the historical top three on the Stratego platform Gravon and was seen at the time as a milestone for AI mastering the game.
Ataraxos takes a different approach. On top of training, it adds decision-time planning: before every move, it runs a search based on its inference of the opponent's hidden pieces, then commits. The researchers say it doesn't guess blindly — it looks across all plausible board states for the safest move.
| Metric | Ataraxos vs. prior work |
|---|---|
| Reinforcement-learning training compute cost | about 1/500 |
| Self-play games | about 1/30 |
| Training samples | about 1/100 |
Beyond the 15-1-4 record against Niemeijer, Ataraxos's overall record against world-championship-caliber top players stands at 39 wins and 2 losses.
From AlphaGo to DeepNash, big labs' game-playing AI results over the past several years have typically come bundled with thousands of chips and months of training, a scale hard for university teams to keep pace with. Ataraxos brings the training bill for this class of problem down to what a single mid-sized research grant can cover, suggesting that search at inference time can substitute for a large share of training-time compute rather than simply adding more self-play. That's the same direction as the broader shift toward moving compute from training to inference that's been playing out across large language models over the past couple of years, where reasoning and search at answer time increasingly do work that used to require ever-bigger pretraining runs.
The paper's authors include CMU's Samuel Sokota (first author) and Zico Kolter; MIT's Gabriele Farina and Zhiyuan Fan; NYU's Eugene Vinitsky; and Stanford's Hengyuan Hu. Whether the code and model weights will be released, and how the system performs on other imperfect-information games, were not addressed in MIT's press release; that will depend on the paper's supplementary materials and any releases from the authors.
Sources: Nature paper, MIT News Office, Tech Xplore, CocoLoop; training configuration, cost basis and match records verified against the paper, with DeepNash comparison figures as stated in the paper.