Jul 2026· CCF Transactions on Pervasive Computing and Interaction· Vol 8, pp. 483 - 499· 0 citations· 42 references
TL;DR
This paper proposes effective multi-agent selective learning methods to boost sample-efficient training by learning from successful experiences, and adopts a retrogression-based selection method to identify successful agent trajectories from the team rewards.
This paper introduces Multi-AGent Preference-Integrated lEarning (MAGPIE), a framework that leverages agent-specific preference signals in the multi-agent learning process and can derive Nash equilibrium solutions.
Ni Mu, Yao Luan, Yiqin Yang et al.· IEEE Transactions on Automat...· 0 citations
This work introduces a novel MARL framework, Multi-Agent Divergence Policy Optimization (MADPO) with Mutual Policy Divergence Maximization (Mutual PDM), and proposes a new extension of CCS divergence for measuring policy divergence of more than two agents, the Generalized Conditional Cauchy-Schwarz (GCCS) divergence.
Haowen Dou, Lujuan Dang, Mingfei Lu et al.· IEEE Transactions on Pattern...· 0 citations
M3W is a novel approach that applies mixture-of-experts (MoE) to world model instead of policy, enabling both learning and planning, and demonstrates superior performance, sample efficiency, and multi-task adaptability.
Zi-Jie Zhao, Zhong Zhao, Kaixuan Xu et al.· Neural Information Processin...· 9 citations
Multi-agent Reinforcement learning has gained significant attention for solving decision-making problems involving multiple autonomous agents. However, effective learning in MARL is still difficult due to environments, dependencies between agents, and poor exploration strategies. Although adaptive exploration and curriculum learning methods, such as Reward Prediction Error Adaptive Learning (RPEAL) along with Reward-Shaped Adaptive Curriculum Learning (RSACL), have produced good outcomes in single-agent reinforcement learning, their use in multi-agent contexts has not been thoroughly investigated. In this research, RPEAL and RSACL are introduced. This paper extends the previous single-agent work to the broader realm of cooperative multi-agent reinforcement learning. The introduced adaptive control mechanism are integrated into several popular multi-agent algorithms such as IPPO, CPPO, MADDPG, and MASAC are empirically compared in a standard petting-zoo environments. The experimental evaluation shows gains in these algorithms upon the introduction of adaptive control mechanisms, where the centralized critic outperforms the individual learners in terms of stability and convergence. Unlike previous works, which only considered single-agent reinforcement learning, in this paper we extend the RPEAL and RSACL to the multi-agent domain. To be specific, we propose team reward prediction error modeling with a centralized critic, as well as performance-driven curriculum learning for multi-agents.
B. Adwaith, Kevin Francis, Remya Nair T· International Conference on...· 0 citations
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents'interaction histories without sacrificing environment sample efficiency remains underexplored. We term this problem Experience Distillation and develop an implementation that requires no further environment interaction beyond the collected experience. Experiments on 749 curated software-engineering tasks and six text-adventure games show that it retains at least 64.8\% of the gains from in-context learning across both domains, whereas direct supervised fine-tuning on the collected experience recovers only 3.8\%. Compared with classical reinforcement-learning baselines, in-context learning from trial-and-error experience followed by Experience Distillation matches their performance with at least \(9.6\times\) fewer environment samples.
As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions that generalize across multiple environments. Because demonstrations in one MDP entangle reward information with that environments specific structure, the resulting rewards frequently fail to generalize when the agent is deployed in a new setting. We first analyze how different feedback modalities constrain rewards, showing that, in the unlimited-data regime, comparisons impose strictly stronger global constraints than other modalities. Beyond this theoretical analysis, we introduce a hierarchical machine teaching algorithm for reward learning that operates across multiple MDPs. The algorithm first greedily selects informative environments that expose complementary reward constraints, then strategically queries low-cost feedback within those environments. Empirically, our method achieves substantially lower regret and stronger generalization to held-out environments than uniform teaching baselines under identical feedback budgets, demonstrating the importance of multi-environment, multi-modal teaching for learning dynamics-robust reward functions.
Ali Larian, Qian Lin, Zong-Wu Chang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.