Skip to content
Book Open access

Combining Policy Gradients with Quality-Diversity in Cooperative Multi-Agent Reinforcement Learning

Jul 2026 · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 0 citations · 17 references

Abstract

Quality-Diversity (QD) methods combined with policy gradients have shown strong performance in single-agent reinforcement learning, but extending them to multi-agent settings introduces challenges from partial observability and agent interactions. We propose MAPGA-ME, a multi-agent extension of PGA-MAP-Elites that integrates policy gradient updates into MAP-Elites for cooperative control. Our results show that directly transferring policy gradient mechanisms from single-agent QD does not consistently improve performance in multi-agent environments. In particular, a design choice effective in single-agent settings becomes less suitable under decentralized, partially observable conditions. Across multiple configurations, we identify key factors affecting the effectiveness of policy gradient-based QD in multi-agent learning, providing practical guidance for adapting these methods.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.