Researchers from Carnegie Mellon, MIT, NYU, and Stanford developed Ataraxos, an AI that succeeded where earlier systems like DeepNash failed, by mastering the imperfect‑information dynamics of Stratego. Trained on a modest cluster of 16 GPUs and a few thousand dollars’ worth of compute, Ataraxos faced Pim Niemeijer, widely regarded as the strongest human player, and recorded a 15‑win‑to‑1‑loss record with four draws. The game’s challenge stems from its hidden‑information structure: each side deploys 40 pieces whose ranks are unknown until a battle reveals the winner’s identity, creating more than a decillion possible initial configurations. Unlike chess, where games typically end after around 40 moves, Stratego can span up to 2,000 moves, and strategic bluffing further complicates decision‑making. Ataraxos leveraged techniques tailored to large‑scale imperfect‑information domains, employing self‑play reinforcement learning combined with policy‑gradient methods to balance exploration of hidden states and long‑horizon planning. The result demonstrates that sophisticated game‑theoretic AI can be built cost‑effectively, turning a historically elusive classic into a solvable challenge.
Read original
arstechnica/ai