Risk-Informed Multi-Agent Reinforcement Learning for Embedded Systems on Resource-Constrained Hardware
Abstract
Learning-enabled control systems increasingly rely on multi-agent reinforcement learning to operate in uncertain and interactive environments. While risk-aware decision-making has been shown to improve safety and robustness, deploying such algorithms on resource-constrained embedded platforms remains a significant challenge due to limited memory, compute, and communication resources. In this paper, we propose a risk-informed multi-agent reinforcement learning framework based on cumulative prospect theory (CPT) and introduce a CPT-SARSA algorithm suitable for real-time execution on low-power embedded hardware. We implement our algorithm on Arduino Nano 33 BLE Sense microcontrollers and evaluate real-time leader-follower coordination in a stochastic grid-world with obstacles and hazards. Our experiments show that CPT-informed agents achieve faster task completion and significantly fewer collisions compared to standard Q-learning, highlighting safety-efficiency tradeoffs under embedded constraints. This demonstration shows that complex multi-agent reinforcement learning algorithms can be deployed on low-cost commodity hardware, increasing access to state-of-the-art learning-enabled control without costly compute resources.