Skip to content

An AI beats the world's best Stratego player

Hidden pieces, bluffing, risk-taking: Stratego resisted machines. A university AI, trained for a few thousand dollars, has just beaten it.

Advertisement

You may have played it as a child. Two armies of forty pieces, placed face down. You know where your bombs and your marshal are. You know nothing about your opponent's.

So you bluff: you advance a simple scout as if it were your strongest piece, hoping to make the enemy retreat.

This board game, Stratego, was one of the last great games where humans still beat machines. That is no longer the case.

Why it was so difficult

In chess or Go, everything is visible. Both players see the same board. An AI can calculate millions of moves, it knows exactly where the game stands.

In Stratego, half the information is hidden. You have to guess what the other player's pieces are from the way they move. You also have to hide your own intentions, and sometimes make them believe the opposite of the truth. The number of possible situations far exceeds that of chess.

In 2022, DeepMind, Google's AI lab, had presented DeepNash, capable of competing with very good online players. But on the sidelines of the 2023 world championship, it had lost against most of the best, including Pim Niemeijer, considered the most decorated player in the game's history.

The result

Researchers from MIT, Carnegie Mellon, NYU and Stanford have just published in the journal Nature the results of their AI, named Ataraxos. Against Pim Niemeijer, it won 15 games, lost only one, and drew 4 matches.

At this level, it is an unprecedented gap. The best players are usually neck and neck, because the game forces you to take risks. Against the best players present on the sidelines of a world championship, Ataraxos racked up 39 wins for 2 losses.

It also beat three multiple world champions in a fast variant of the game, a first, and set records in other games with hidden information, such as the cooperative card game Hanabi.

The detail that changes everything: the price

The most impressive thing may not be the victory. It is what it cost. According to the scientific paper, training Ataraxos for Stratego cost a few thousand dollars, with less than a hundredth of the training examples used for DeepNash.

A university budget succeeded where the millions of a tech giant had failed. It is good news for public research, and a reminder that money does not buy everything.

How it plays

Ataraxos learned by playing against itself, again and again, keeping what worked. During the game, it imagines what the opponent's hidden pieces might be, and plans accordingly.

And above all, it is unpredictable. The researchers stress this: in Stratego, a player who is too consistent is quickly unmasked and exploited. To win, you sometimes have to leave an element of chance in your choices. Exactly like a good poker player.

A fitting name 🧘
Ataraxos comes from the Greek "ataraxia", a word dear to ancient philosophers such as Epicurus: the absence of trouble, the tranquillity of the soul. The researchers say they chose it to designate someone whom nothing worries. For a bluffing player, it is ideal: never let anything show.

And outside the game?

The researchers envisage applications wherever information is hidden: business negotiations, cybersecurity, or even military manoeuvres.

Caution is needed. A game has fixed rules, known to all. The real world does not. But the question this victory raises is very real, and it is the subject of our reflection this evening: is bluffing lying?

Advertisement