Skip to content
Preprint

Emergence of cooperation: A reputation-modulated reinforcement learning

Aug 2026 · 0 citations · 4 references
Physics Biology Mathematics

TL;DR

The results reveal that reputation-modulated learning significantly promotes the emergence of cooperative behavior, and the discontinuous phase transition from full cooperation to full defection as the temptation increases.

Abstract

Reputation is widely recognized as a key mechanism for sustaining cooperation. However, most existing game-theoretic models treat reputation primarily as an external factor that modulates payoffs, interaction structures, or strategy update rules. In many social contexts, though, reputation operates primarily as information -- it shapes how individuals interpret their own experiences and assess the behavior of others. To bridge this gap, we propose a spatial prisoner's dilemma game grounded in the reinforcement learning paradigm, in which agents equipped with Q-learning integrate both individual and social information via a locally defined reputation metric to guide their decisions. Our results reveal that reputation-modulated learning significantly promotes the emergence of cooperative behavior, and we observe a discontinuous phase transition from full cooperation to full defection as the temptation increases. Cooperation spreads through the nucleation of cooperative clusters, whereas the disintegration of these clusters drives the system into an absorbing state of complete defection. Overall, this study demonstrates that reputation facilitates cooperation not only by providing direct incentives but also by reshaping the social information landscape that agents rely on for learning and adaptation.

View source

Similar papers

Preprint Sep 2026

Learning to cooperate in a changing world: How caring about the future promotes cooperation across scales

In social dilemmas, individuals need to forgo short-term temptations to achieve synergistic collective outcomes through cooperation. Previous work has examined mechanisms through which cooperation can evolve, including direct reciprocity, indirect reciprocity, environmental stochasticity, network reciprocity, and demog...

Yu-Xin Geng, Xing-Ru Chen, Xin Wang et al. · 0 citations
Preprint Aug 2026

Evolution of cooperation with Q-learning: how much information do we need?

Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability, which allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload.

Yi-Hsin Ku, Xin Ou, Ji-Qiang Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Evolutionary Stability Does Not Guarantee Learning Accessibility: A Multi-Agent Reinforcement Learning Perspective on Cooperation Emergence

Cooperation emergence is a central problem in multi-agent systems because decentralized agents must coordinate while adapting to the changing behavior of others. Evolutionary game theory identifies strategically stable outcomes, but stability under a population adjustment dynamic need not imply that finite-sample learn...

Yi-Jie Wang · 0 citations
Preprint Sep 2026

Payoffs and perception mediate environmental feedback in an N-player trust game with Q-learning

Trust develops through learning, while collective behavior can alter the environment in which later decisions are made. Reinforcement-learning models describe adaptation, and eco-evolutionary models describe behavior-environment feedback, but how an endogenous environment changes trust through material incentives and p...

Ru-Qiang Guo, Zhao-Yi Hu, Fang Wang et al. · 0 citations
Preprint Aug 2026

Emergence of Reputation-Based Cooperation in LLM Agents

Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language prompts, evolve through cultural transmiss...

Kazuya Horibe, Kenji Itao, Wataru Toyokawa · 0 citations
Open access Sep 2026

Research on the Game Dynamics of Optional Public Goods

The main findings demonstrate that individuals update strategies primarily through self-adjustment based on historical payoffs, with imitation playing merely an auxiliary role and the optimal self-adjustment proportion is approximately 0.9, and low sensitivity coefficients and mutation rates favor the emergence of coop...

Hao-Chen Wu, Meng-Cheng Sun, Lu-He Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.