Reinforcement Learning-Driven Multi-Agent Artificial Intelligence for Autonomous Enterprise Process Optimization
Abstract
The more complex and interdependent nature of enterprise processes like supply-chain management, procurement and scheduling means static rule-based automation is becoming increasingly ineffective. This paper suggests a Reinforcement Learning-Driven Multi-Agent Artificial Intelligence framework for autonomous enterprise process optimization, where the agents are specialized to observe local process states, communicate with one another via a shared communication layer and collectively optimize the performance of the system without the involvement of a centralized controller. The problem is represented as a Decentralized Partially Observable Markov Decision Process and agents are trained in a centralized-training, decentralized-execution setting with a hybrid local/ global reward structure. The proposed multi-agent framework reduces the cycle-time by 32.6 %, is 27.8 % cost efficient and improves throughput by 25.3% compared to a single-agent RL (21.4%, 18.6%, 16.9%) and a static optimization solver (14.7 %, 12.3 %, 10.8 %), while achieving convergence in 2,650 episodes against 4,200 episodes for a single-agent RL framework. Scalability analysis between 4, 8, 16 and 32 agents reveal a graceful reduction in cycle-time (from 34.1% to 25.3%), and a corresponding increase in coordination overhead (from 12 to 121 ms/step), which define the practical operating range of 8-16 agents. The results show that the coordinated multi-agent reinforcement learning method is feasible for autonomous enterprise process optimization.