Hierarchical Reinforcement Learning Control of a Quadruple Tank Plant Under Partial Observability
Reinforcement learning (RL) has emerged as a promising paradigm for controlling nonlinear and highly coupled dynamical systems, particularly in scenarios where accurate models are difficult to obtain. This paper investigates architectural strategies for deploying deep reinforcement learning controllers in strongly coupled Systems of Systems under partial observability, using the Quadruple Tank Plant (QTP) as a benchmark case study. A centralized Deep Deterministic Policy Gradient controller is first trained with full state information and evaluated in simulation and on a real laboratory-scale QTP. The control problem is then decomposed into decentralized agents with limited local observations, which independently generate control proposals. To coordinate these decentralized decisions, a supervisory reinforcement learning agent is introduced, learning to fuse and minimally correct the agents’ proposals based on partial system measurements. The proposed hierarchical architecture enables effective control even in scenarios where subsystem objectives are locally feasible but globally conflicting. Experimental results in simulation and on real hardware demonstrate that the decentralized–supervised architecture can achieve comparable or improved aggregate tracking performance relative to a centralized policy, while preserving decentralized proposal generation and enabling execution-time supervisory coordination under partial observability. The results also show stable sim-to-real behavior in the presence of sensor noise, actuator delays, and modeling inaccuracies.