Skip to content
Conference

Subjective Decision-Making in Multi-Agent Reinforcement Learning with Cumulative Prospect Theory and Social Value Orientation

Aug 2026 · Conference on Control Technology and Applications · pp. 712-718 · 0 citations · 29 references

Abstract

Cyber physical systems such as autonomous vehicles operate in highly dynamic environments where interactions with other autonomous and human agents is inevitable. Reinforcement learning (RL) is a well-established paradigm to allow agents to learn behaviors through interactions with the environment when a model of the environment is not known or available. When actions of any single agent in such multi-agent setups can be influenced by its own belief on other agents’ actions and their risk-taking abilities, it becomes critical to quantify social preferences for each agent. In such a situation, it is also important to model widely observed preferences of human operators who will likely be sharing the same environment as autonomous agents. While significant progress has been made in modeling social preferences and risk-aware behavioral models into RL, these have largely occurred in parallel. This paper presents a two-pronged solution approach that uses cumulative prospect theory (CPT) to characterize risk-awareness of an individual agent and social value orientation (SVO) to represent the spectrum of egoistic to altruistic agent behaviors. We define the SVO-informed CPT-value of a random variable, and use it to design a novel subjective learning procedure that we term the CPT-SVO-MARL Algorithm. We empirically demonstrate stable learning behavior in a representative cyber physical control scenario.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.