Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play
Abstract
Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.