STAGE: Spatio-Temporal Aggregation via Graph Embedding for Multi-Agent Reinforcement Learning in Industrial Optimization
Abstract
Industrial multi-agent coordination requires distributed subsystems to collaborate under heterogeneous relationship structures whose relative importance shifts across operational contexts—physical constraints dominate startup while operational hierarchies govern steady-state. Existing multi-agent reinforcement learning approaches either ignore these structural distinctions or aggregate them uniformly, limiting adaptive coordination capabilities. This paper presents STAGE (Spatio-Temporal Aggregation via Graph Embedding), integrating multi-layer graph processing with spatio-temporal learning for context-dependent coordination. The architecture processes distinct relationship types through dedicated attention mechanisms with learned adaptive fusion, enabling coordination emphasis to adjust dynamically across operational phases. Spatio-temporal integration couples multi-layer spatial structures with multi-scale temporal dynamics through attention-based fusion mechanisms, while graph-guided hypernetworks generate mixing weights that preserve the Individual-Global-Max property essential for decentralized execution. Comprehensive evaluation on steam power plant coordination demonstrates that STAGE significantly outperforms existing multi-agent and optimization methods in both learning efficiency and asymptotic performance, while providing interpretable coordination mechanisms and maintaining the monotonicity property critical for decentralized industrial deployment. Note to Practitioners—Industrial subsystems interact through multiple relationship types whose importance varies across operational contexts. We propose a framework that separately processes these relationships and learns to adjust their emphasis adaptively. Graph-guided coordination ensures individual actions optimize system-wide performance. Evaluation on power plant configuration demonstrates superior results over conventional optimization and learning methods. The approach learns from operational data while respecting the simulator’s physical feasibility constraints. We believe this framework has strong potential for chemical processing, manufacturing, and energy management systems.