Skip to content
Preprint

Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination

Aug 2026 · 1 citation · 47 references
Computer Science

TL;DR

BayesBeliefAgent is introduced, which pairs a hierarchical LLM planner with a Bayesian tracking module and evaluates performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill.

Abstract

Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows that a partner has changed its skill. Existing methods either treat partner tracking as passive context-leaving the agent aware of the shift but slow to act-or replan indiscriminately. We introduce BayesBeliefAgent, which pairs a hierarchical LLM planner with a Bayesian tracking module. Rather than replanning constantly, our agent interrupts its current skill only when a partner's actions directly contradict the inferred skill. Beyond standard reward, we evaluate performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill. Across benchmark Overcooked environments, contradiction-conditioned control drastically narrows this belief-action gap while requiring an order of magnitude fewer replans than heuristic methods

View source

Similar papers

Preprint Aug 2026

Training Small LLMs as Spatial Multi-Agent Policies

Training LLM-based multi-agent systems with multi-agent reinforcement learning with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.

Yi Mao, Andrew Perrault · 0 citations
Preprint Aug 2026

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

This work proposes AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning that aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space.

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao et al. · 1 citation
Preprint Aug 2026

History Matters: Meta-policy Delegation with Heterogeneous Multi-agent Reinforcement Learning

This paper develops a multi-agent reinforcement learning-based (MARL) delegation training that enables agents to make sequential delegation decisions while minimizing the total execution cost and introduces two new frameworks for collaboration and delegation in multi-agent systems.

Ziqing Lu, Avinash Mudireddy, Sarra M. Alqahtani et al. · 0 citations
Preprint Aug 2026

The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate

The collaboration tax is formulated as the team-decentralisation loss of a two-player cooperative game with private information, with two propositions characterising its sign and its equivalence to a max-superadditivity violation.

Wei-Xiang Sun, Zehong Wang, Hong Huang et al. · 0 citations
Jul 2026

Multi-Agent LLMs Fail to Explore Each Other

This work introduces Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection that substantially improves exploration behavior and downstream task performance and shows theoretically that the value of exploration increases with agent diversity.

Hyeong Kyu Choi, Jiatong Li, Wendi Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.