Skip to content
Open access

Multi-Agent Reinforcement Learning via Agent-Specific Preference

Aug 2026 · IEEE Transactions on Automation Science and Engineering · Vol 23, pp. 14725-14741 · 0 citations · 58 references
Computer Science

TL;DR

This paper introduces Multi-AGent Preference-Integrated lEarning (MAGPIE), a framework that leverages agent-specific preference signals in the multi-agent learning process and can derive Nash equilibrium solutions.

Abstract

Multi-agent reinforcement learning (MARL) is a powerful framework for solving complex collaborative tasks, but it relies heavily on well-defined global reward functions. Designing such rewards is challenging, especially in systems with heterogeneous agents, where a single scalar objective may fail to capture diverse behaviors. In this paper, we introduce Multi-AGent Preference-Integrated lEarning (MAGPIE), which addresses these challenges through agent-specific preference modeling. Each agent is evaluated by a dedicated expert through preference signals, eliminating the need for global evaluation. We theoretically prove that optimizing these decentralized pReferences converges to a Nash equilibrium policy. To integrate local preferences into a coherent global objective, we construct agent-specific reward models from preference data and combine them via a monotonic aggregation mechanism. We further prove that optimizing this aggregate reward model is equivalent to training the Nash equilibrium policy. Extensive experiments on benchmark multi-agent tasks and a sequential production line task show that MAGPIE achieves performance comparable to reward-engineered baselines, demonstrating its potential to facilitate policy learning in scenarios where precise reward engineering is impractical. Note to Practitioners—Multi-agent systems are widely used in modern engineering applications. For example, autonomous vehicle fleets coordinate to prevent collisions while maintaining efficiency, and industrial manufacturing lines work together to meet production targets without causing buffer overflows. Multi-agent reinforcement learning (MARL) provides a powerful framework for enabling such collaboration, but its success depends heavily on well-designed reward functions. Designing these rewards is often challenging, especially when agents play distinct roles, as it is difficult to translate complex interactions and diverse agent behaviors into precise numerical signals. In contrast, providing comparative feedback on preferred behaviors is often more intuitive than specifying explicit mathematical rewards. In this paper, we introduce Multi-AGent Preference-Integrated lEarning (MAGPIE), a framework that leverages agent-specific preference signals in the multi-agent learning process. MAGPIE learns agent-specific reward models and combines them into a unified global objective using monotonic aggregation. By optimizing this objective, we can derive Nash equilibrium solutions. Importantly, preferences can be provided by lightweight automated rules or domain-specific heuristics, eliminating the need for costly human annotators. MAGPIE is effective, easy to implement, and particularly suitable for complex systems where traditional reward design is impractical.

Read PDF

Similar papers

Open access Jul 2026

Sample-Efficient Multi-Task and Multi-Objective Reinforcement Learning by Combining Multiple Behaviors

One of the main challenges in the field of artificial intelligence, and reinforcement learning (RL) in particular, is the development of generalist and flexible agents capable of solving multiple tasks—each requiring the agent to learn a potentially new, specialized behavior. Tackling this challenge requires agents to learn behaviors that may involve optimizing a single objective, or trading off between multiple conflicting objectives. In this thesis, we study how to design flexible RL agents that can, in a sample-efficient manner, adapt their behavior to solve any given tasks—each of which is defined by multiple (possibly conflicting) objectives. We introduce new multi-policy methods that empower RL agents to (i) carefully learn multiple behaviors, each specialized in a particular task; and (ii) combine previously-learned behaviors to efficiently identify solutions to novel tasks, which, importantly, may require the agent to assign different preferences to each of its new objectives. The methods we introduce have strong theoretical guarantees regarding the optimality of the set of behaviors learned by agents and their capability to solve new tasks in a zero-shot manner, even in the presence of function approximation errors. We evaluate the proposed methods in various challenging multi-task and multi-objective RL problems and show that our algorithms outperform various current state-of-the-art methods in domains with both discrete and continuous state and action spaces.

L. N. Alegre, Ana L. C. Bazzan, Bruno C. da Silva · 0 citations
Preprint Aug 2026

History Matters: Meta-policy Delegation with Heterogeneous Multi-agent Reinforcement Learning

This paper develops a multi-agent reinforcement learning-based (MARL) delegation training that enables agents to make sequential delegation decisions while minimizing the total execution cost and introduces two new frameworks for collaboration and delegation in multi-agent systems.

Ziqing Lu, Avinash Mudireddy, Sarra M. Alqahtani et al. · 0 citations
Jul 2026

Reinforcement Learning: From Algorithms To Foundation Models

This thesis develops diffusion-based world models, investigates RL for efficient video generation, explores generative models as policy classes, and studies interactive video world models in which actions shape future observations, and addresses long-horizon modeling through architectures with memory.

Zihan Ding · 0 citations
Jul 2026

Efficient Heterogeneous Exploration with Mutual Policy Divergence Maximization for Multiagent Reinforcement Learning.

This work introduces a novel MARL framework, Multi-Agent Divergence Policy Optimization (MADPO) with Mutual Policy Divergence Maximization (Mutual PDM), and proposes a new extension of CCS divergence for measuring policy divergence of more than two agents, the Generalized Conditional Cauchy-Schwarz (GCCS) divergence.

Haowen Dou, Lujuan Dang, Mingfei Lu et al. · 0 citations
Book Open access Jul 2026

Combining Policy Gradients with Quality-Diversity in Cooperative Multi-Agent Reinforcement Learning

Quality-Diversity (QD) methods combined with policy gradients have shown strong performance in single-agent reinforcement learning, but extending them to multi-agent settings introduces challenges from partial observability and agent interactions. We propose MAPGA-ME, a multi-agent extension of PGA-MAP-Elites that integrates policy gradient updates into MAP-Elites for cooperative control. Our results show that directly transferring policy gradient mechanisms from single-agent QD does not consistently improve performance in multi-agent environments. In particular, a design choice effective in single-agent settings becomes less suitable under decentralized, partially observable conditions. Across multiple configurations, we identify key factors affecting the effectiveness of policy gradient-based QD in multi-agent learning, providing practical guidance for adapting these methods.

Hai D. Pham, Ngoc Hoang Luong · 0 citations
2025

Learning and Planning Multi-Agent Tasks via an MoE-based World Model

M3W is a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning, and demonstrates superior performance, sample efficiency, and multi-task adaptability.

Zi-Jie Zhao, Zhong Zhao, Kaixuan Xu et al. · 9 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.