← All articles

The Ataraxos AI beat the best Stratego player in history — and learned for thousands of dollars, not millions

A floating light glass game board with rows of smooth pieces, some glowing teal and coral

Ataraxos, an AI built by researchers from Carnegie Mellon, MIT, New York University and Stanford, has beaten the best Stratego player in history: 15 wins, 1 loss and 4 draws in 20 games against Pim Niemeijer, the game's most decorated player. The result is published in Nature. The striking part is not the win itself but its price: training took 16 GPUs for a week plus 4 GPUs for four days — a few thousand dollars — while an earlier DeepMind attempt needed roughly 3–4.5 million dollars of compute. The Stratego story is really a story about making decisions when half the situation is hidden, which is exactly the problem in negotiations, financial markets and war games.

What Stratego is and why it resisted AI

In Stratego each player has 40 pieces representing military ranks, from marshal down to spy, plus bombs and a flag. Your opponent sees where your pieces are but not what they are. Identities are revealed only on collision: the weaker piece is removed and the winner's identity becomes known. You win by capturing the enemy flag.

Hidden information is what makes Stratego hard for machines. In Texas Hold'em a player has just two hidden cards — 1,326 possible hands, few enough for a computer to enumerate. In Stratego, 40 pieces can be arranged in any order: more than 10^33 setups. Add game length — chess rarely exceeds 40 moves, while Stratego easily runs to 2,000 — and you get a problem where brute-force search simply does not work.

There is a third difficulty: bluffing. Sometimes it pays to move a weak piece as if it were a marshal to scare your opponent. Bluff too often and your threats mean nothing; never bluff and you become predictable. That balance is what stumped earlier AIs. Even DeepMind's DeepNash, introduced in 2022, never reliably beat the strongest humans despite a huge budget.

The two ideas that changed the game

Like DeepNash, Ataraxos learned by playing against itself — 163 million games in total. Winning moves were reinforced and played more often, losing moves less. But hidden information sends such algorithms around in circles, because the value estimate keeps shifting. The team handled this by changing strategy unevenly: bold, large steps early in training and small, careful ones later.

The bigger innovation was something DeepNash never had: thinking ahead before each move. In AlphaGo, a search refined the general strategy, but Stratego's search space is so large that it was unclear whether that could work at all. The answer is a second neural network, a belief model. It watches how the opponent moves and estimates which pieces might occupy the hidden cells. Ataraxos no longer iterates over every arrangement — it samples plausible ones, plays candidate moves in them, and decides based on the outcome.

A floating light glass game board with rows of smooth pieces, some glowing teal and coral

Calm and unbothered

The name Ataraxos comes from the ancient Greek word for calm and unbothered. And it shows in its play. Where a human facing a rout takes a gamble, the AI works its way back methodically. Knowing a secret does not distract a machine the way it distracts a person — a human finds it hard to make decisions while ignoring known hidden information, a machine finds it easy. That is why Ataraxos can "bluff" with a straight face and come back from positions where it had only about a 2% chance of winning.

Niemeijer is a four-time world champion and spent more than 600 weeks ranked number one in the world. Over three weeks he played 20 games against the AI, earning $100 for each win, and won exactly once. The researchers see no flaw in that loss: good Stratego demands a randomized setup, so luck always plays a role. At the 2025 Stratego World Championship, challengers fared even worse — the AI won 38 of 40 games.

The numbers: 16 GPUs versus 1,024 chips

DeepNash trained for two to three months on 1,024 of Google's specialized chips — the Ataraxos team estimates that at $3–4.5 million in 2025 prices. Ataraxos needed 16 GPUs for a week plus 4 GPUs for four days to train its belief model. It played about 34 times fewer games and still ended up much stronger. The secret is a custom simulator that runs millions of moves per second on a graphics card.

Not just Stratego

The same approach worked in other hidden-information games. In Barrage Stratego, a faster eight-piece variant, the AI beat three world champions. In the cooperative card game Hanabi and the Chinese card game dou dizhu, it reached the level of the best bots. This is no longer a one-off solution to a single game but a general design pattern for reinforcement learning and search under large amounts of hidden information.

The researchers are looking beyond board games. Real problems — negotiations, markets, military scenarios — have no fixed rules and no clear winner, but the team notes that any analysis of a real situation starts with a simplified model. For rehearsing responses to a strong opponent, the approach already works. The caveat is honest too: Ataraxos cannot yet explain why it makes the moves it makes, and that is the team's next task.

What it means in practice

The main takeaway is simple: what decides is not the size of the budget but the efficiency of the method. Ataraxos spent orders of magnitude less compute than its predecessor and still played stronger. Generative AI follows the same logic: at volume, the winner is not the heaviest model but the one that is cheaper and more reliable for the job. According to our service's data over the last 30 days (window 05.09–05.10.2026, NeuralSpace production database), the most in-demand image model was by no means the most expensive: nano-banana-2 accounted for 2,351 generations at an average cost of roughly $0.053 per image. Video looks the same: the lightweight grok-image-to-video logged 1,648 tasks at $0.15 each, while the heavy seedance-2.5 managed only 246 at an average of $1.74. In total over 30 days the platform completed 12,689 generations: 9,907 images, 2,334 videos and 448 music tracks (anonymized aggregate, server/seo/benchmarkData.json).

FAQ

What is Ataraxos?

A neural network for Stratego built by a team from CMU, MIT, NYU and Stanford. It is the first in history to beat the best human player, and it does so with incomparably less compute than earlier attempts.

How is this different from AlphaGo beating Go?

In Go and chess both players see the whole position. Stratego is a game of hidden information: half of the opponent's pieces are unknown. That is why it required not only reinforcement learning but also a model that estimates hidden pieces, plus search over plausible arrangements.

Can I try models like this?

Ataraxos itself is a research project with no public access. But you can talk to modern language models in NeuralSpace Chat, and build your own AI project in Code.