Autonomous Data Pipelines Using Deep Reinforcement Learning for Self-Optimization in Cloud Environments
Abstract
The growth of cloud-based data infrastructure has skyrocketed and the complexities in managing and optimizing data pipelines at scale has never been higher. Conventional pipeline orchestration techniques are based on models that use rigid configurations and rule-based scheduling systems, which are inherently non-adaptive to the variations in workloads, resource heterogeneity, and data evolution inherent to modern cloud settings. This work provides a new framework for autonomous self-optimization of data pipelines with Deep Reinforcement Learning (DRL), where intelligent agents continuously observe the execution states of the pipeline, learn the best methods for resource provisioning and autonomously reconfigure individual components to maximize throughput while minimizing latency and cost. The proposed framework utilizes a hybrid DRL architecture by integrating an approach called DQN which is based on deep learning, to handle discrete scheduling decisions including resource provisioning decisions and an approximate actor-critic method, i.e. PPO, that operates on continuous control problems for capturing efficiency in running the workloads simultaneously. We formulate a multi-objective reward function to trade off multiple competing optimization objectives of maximizing execution efficiency and minimizing cost, while achieving the desired level of fault tolerance over heterogeneous cloud infrastructure. Experimental evaluations over simulated and real-world cloud environments show up to 68.3% pipeline throughput improvement, 71.2% end-to-end latency reduction, and operational cost saving of 54.7%, compared with traditional static or threshold-based pipeline management systems. This new framework turns out to be a solid basis for building fully autonomous, self-healing and cost-effective cloud data infrastructure.