DeepMind details Ataraxos, an AI system built for Stratego's hidden-information gameplay
Researchers describe Ataraxos, a decision-making system for the imperfect-information board game Stratego. It combines two interdependent self-play reinforcement learning processes—one for choosing piece set-ups, one for choosing moves—built on transformer networks, alongside a belief network trained on self-play games and a search procedure used at test time.
GoKawiil's interpretation of the reporting above, not reported fact.
Games with hidden information, like Stratego, are widely regarded as harder testbeds for AI than perfect-information games such as chess or Go, since agents must reason about opponents' unseen states rather than just optimize known positions. The researchers' choice to split set-up and move learning into separate specialized processes, rather than one end-to-end model, suggests that decomposing complex decision problems by phase can outperform unified architectures, an approach that could inform how AI systems are designed for other multi-stage, uncertain real-world tasks.
- Ataraxos uses two interdependent self-play RL processes for Stratego's set-up and move phases.
- A belief network trained on self-play games helps handle the game's hidden information.
- Set-up learning uses a decoder-only architecture while move learning uses an encoder-only one, reflecting different optimal training methods for each phase.
Source: nature.com — Sokota, 2026-09-30
Published there as: “Scalable decision-making for games of imperfect information”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.