Résumé
This paper studies the last-iterate convergence properties of the exponential weights algorithm with constant learning rates. We consider a repeated interaction in discrete time, where each player uses an exponential weights algorithm characterized by an initial mixed action and a fixed learning rate, so that the mixed action profile pt played at stage t follows an homogeneous Markov chain. At first, we show that the probability to play a strict Nash equilibrium at the next stage converges almost surely to 0 or 1. Secondly, we show that the limit of pt, whenever it exists, belongs to the set of Nash Equilibria with Equalizing Payoffs. Finally, we show that in strong coordination games, where the payoff of a player is positive on the diagonal and 0 elsewhere, pt converges almost surely to one of the strict Nash equilibria.
Référence
Maurizio d'Andrea, Fabien Gensbittel et Jérôme Renault, « Games played by exponential weights algorithms », Mathematics of Operations Research, 2026, à paraître.
Publié dans
Mathematics of Operations Research, 2026, à paraître
