• Home  
  • AI defeated a four-time Stratego world champion 15 to 1 using far less training
- Technology

AI defeated a four-time Stratego world champion 15 to 1 using far less training

A new Nature study introduces Ataraxos, an AI system that mastered strategic decision-making under hidden information, defeating a four-time Stratego world champion 15 to 1 while using dramatically less training than earlier systems.

Editorial illustration of an artificial intelligence system analysing a hidden-information strategy board game

Artificial intelligence has surpassed elite human performance in chess, Go and a growing collection of games. Yet games in which crucial information remains hidden have proved substantially harder. A player cannot simply calculate from a complete board when the opponent knows things the machine does not.

A new study published in Nature on 30 September 2026 reports a major step across that divide. Researchers developed Ataraxos, an artificial intelligence system designed to learn and search in games with imperfect information. In a direct 20-game evaluation, it defeated four-time Stratego world champion Pim Niemeijer with 15 wins, one loss and four draws. That corresponds to an 85% effective win rate when draws count as half a win.

The result matters beyond a board game. Many consequential decisions involve hidden information, from negotiations and markets to logistics and strategic planning. The researchers argue that the methods behind Ataraxos provide a more general design pattern for decision-making when an agent must act without knowing the full state of the world.

Why Stratego remained difficult for AI

Stratego presents a different computational problem from games such as chess. Each player begins by privately arranging 40 pieces. The identities of opposing pieces remain hidden until interactions reveal them, meaning that strong play requires reasoning about what the opponent may possess and how likely different hidden configurations are.

That uncertainty makes conventional search difficult. An AI cannot evaluate only the visible board because many possible underlying states could produce the same observations. It must maintain beliefs about those possibilities while also choosing actions that exploit useful information without exposing too much of its own.

Previous industrial efforts had already demonstrated that reinforcement learning could produce powerful Stratego systems, but the authors note that top-human-level performance under massive hidden information remained costly and difficult. Ataraxos was designed around a combination of efficient self-play, explicit modelling of hidden information and search performed at decision time.

Learning a strategy by playing itself

The system begins with tabula rasa self-play reinforcement learning. Rather than learning directly from human games, it repeatedly plays against itself and shifts its policy towards decisions associated with better outcomes.

The researchers separated Stratego into two interdependent learning processes. One network learns how to arrange the 40 pieces at the beginning of a game, while another learns how to move them once play begins. The processes remain connected because the arrangements produced by the first network become the boards on which the second network plays, and the eventual game outcomes inform both.

A central technical challenge is keeping self-play stable. In imperfect-information environments, learning can become cyclical or unstable because a strategy that exploits one opponent can itself become exploitable by another. Ataraxos coordinates the strength of regularisation with the size of policy updates. Stronger regularisation and larger updates are used earlier, with both adjusted as training progresses. The aim is to learn rapidly without collapsing into a narrow strategy that can no longer improve.

The researchers also developed a highly efficient simulator capable of executing millions of state updates per second on graphics-processing hardware. According to the study, the complete Stratego system cost only a few thousand US dollars to train. The accompanying analysis reports that it achieved stronger playing performance than the earlier DeepNash system while using less than one hundredth as many training examples and less than one thirtieth as many self-play games.

The AI does not simply guess what it cannot see

Self-play provides Ataraxos with a strong underlying policy, but the system adds another component for hidden information. A belief network is trained on the final self-play games to estimate the identities of pieces that remain concealed.

During self-play, the true identities of hidden pieces are available to the training system even though they would not be visible to a player during an actual match. The belief network therefore learns to associate the information a player can observe with probability distributions over the concealed state.

Before making a move, Ataraxos samples plausible hidden game states from this belief model. It then evaluates candidate actions across those possible states using its policy and value network. The resulting estimates support an additional policy update applied only to the current decision. This allows the AI to adapt its immediate choice to the particular information available in the game without retraining its overall strategy against the opponent.

A decisive test against an elite human player

The researchers evaluated Ataraxos against Pim Niemeijer, whose record includes four world championships, 15 Dutch national championships, two online world championships and more than 600 weeks as the number-one ranked player.

The evaluation comprised 20 games spread across three weeks. This structure reduced fatigue and gave Niemeijer time to study the AI between games. Importantly, the human player could adapt to patterns he observed in Ataraxos, while the AI itself did not learn from or adapt its underlying model to him during the series.

Ataraxos won 15 games, lost one and drew four. Counting a draw as half a victory produces an effective win rate of 85%. The researchers describe the margin as unprecedented at the highest level of Stratego. Under a simplified assumption that the game outcomes were independent and identically distributed, an exact one-sided binomial test of whether Ataraxos was more likely to win than lose would produce a P value below 2.6 × 10−4. The authors caution, however, that repeated games against an adapting human opponent do not truly satisfy that statistical assumption.

The system was subsequently demonstrated at the 2025 Stratego World Championship. Across 40 games offered to championship attendees, Ataraxos recorded 38 wins and two losses, equivalent to a 95% effective win rate. This provided a broader test against different human playing styles rather than a single elite opponent.

The method transferred beyond one game

A system that only mastered Stratego would still be an impressive specialist. The researchers therefore tested whether the same underlying design could work across different forms of imperfect information.

They adapted Ataraxos to Barrage Stratego, the cooperative card game Hanabi and the Chinese card game dou dizhu. These environments differ substantially. Barrage remains adversarial but uses fewer pieces, Hanabi requires players to cooperate while holding incomplete information, and dou dizhu combines hidden cards with asymmetric team play.

In Barrage Stratego, the AI was evaluated against three multi-time world champions and achieved superhuman performance. In Hanabi, the system established state-of-the-art results across variants involving two to five players. With search used by all players, average scores ranged from 24.410 to 24.863 out of a maximum of 25, depending on the number of players. The proportion of perfect games reached 89.90% in the three-player setting and 88.40% with four players.

For dou dizhu, Ataraxos also outperformed existing systems. Its role-averaged score was 0.199 against PerfectDou and 0.350 against DouZero, while the underlying policy without search produced smaller advantages of 0.152 and 0.286 respectively. That comparison is important because it isolates some of the additional benefit produced by decision-time search.

Efficiency may be as important as the win record

The headline result is the victory over an elite human, but the computational efficiency may prove more consequential for future research. High-performing game systems have sometimes depended on enormous training budgets, limiting who can reproduce or extend them.

Ataraxos demonstrates that carefully structured learning and search can reduce those requirements substantially. The researchers report total Stratego training costs of only a few thousand dollars. In benchmark matches against five established Stratego bots, Ataraxos won between 97 and 99 of 100 games against each opponent, suggesting that the lower training cost did not come at the expense of playing strength.

The broader implication is not that board-game algorithms can be transferred directly into high-stakes real-world systems. Instead, the work provides evidence that reinforcement learning and search can remain effective when large amounts of relevant information are hidden, provided that uncertainty itself is modelled as part of the decision process.

What the study cannot yet establish

There are important limits to that interpretation. Games provide precisely defined rules, objectives and outcomes. Real decisions often involve changing rules, ambiguous goals, unreliable observations and consequences that cannot be represented by a simple win or loss.

The flagship human comparison also involved only one opponent across 20 games. Niemeijer’s exceptional record makes him a demanding benchmark, and the additional world-championship games broaden the evidence, but they do not represent every possible human strategy. The authors further acknowledge that the repeated matches were not statistically independent because the human player could adapt over time.

Interpretability remains another challenge. A system may calculate risk effectively without being able to explain its reasoning in terms that a human decision-maker can audit. The researchers identify interpretable decision-making as an important direction for future development, particularly before methods inspired by these systems could be considered for consequential applications.

A broader test for decision-making under uncertainty

The significance of Ataraxos lies in the combination of three capabilities: learning strong strategies through self-play, constructing probabilistic beliefs about information it cannot observe, and using additional computation at decision time to refine an action.

That combination allowed the system to dominate one of the strongest Stratego players in history while using far less training than earlier approaches. More importantly, its performance across adversarial, cooperative and team games suggests that the underlying approach is not tied to a single set of rules.

For artificial intelligence research, hidden information has long marked a boundary between environments that can be searched relatively directly and those in which an agent must reason about what it does not know. This study provides strong evidence that the boundary is becoming more tractable, although moving from games to messy real-world decisions remains a much larger challenge.

Source Information

Study: Scalable decision-making for games of imperfect information

Authors: Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter, and colleagues, including Gabriele Farina

Journal: Nature, volume 658, pages 55–59 (2026)

Published: 30 September 2026

DOI: 10.1038/s41586-026-11036-y

Study type: Self-play reinforcement-learning and test-time-search development with human and computational benchmark evaluations across imperfect-information games

Contact Us

Research Today is a South African digital publication that makes credible research easier to understand.

TERMS OF USE & PRIVACY POLICY

follow us