Skip to content
Open access

Intelligent Workflow Scheduling for Distributed Data Processing Systems

2024 · International Journal of Data Engineering and Intelligent Computing · 0 citations

TL;DR

An intelligent workflow scheduling framework that improves performance through adaptive decision-making, predictive analytics, and machine learning, and addresses key challenges like load balancing, scalability, energy efficiency, and fault tolerance is proposed.

Abstract

The rapid growth of data-intensive applications in scientific computing, enterprise analytics, and cloud services has increased the demand for efficient distributed data processing systems. Traditional scheduling methods like FCFS, Round Robin, and heuristic approaches often fail to meet the dynamic and heterogeneous requirements of modern environments. This paper proposes an intelligent workflow scheduling framework that improves performance through adaptive decision-making, predictive analytics, and machine learning. The system dynamically allocates tasks based on resource availability, workflow dependencies, and historical execution data, enabling it to anticipate bottlenecks and reassign tasks proactively. It also incorporates resource heterogeneity modeling and dependency-aware scheduling to reduce idle time and optimize execution. Performance is evaluated using metrics such as makespan, throughput, resource utilization, and fault tolerance, showing significant improvements over traditional methods. The framework also addresses key challenges like load balancing, scalability, energy efficiency, and fault tolerance. Overall, the proposed approach enhances system efficiency and scalability while supporting integration with emerging technologies such as edge computing and hybrid cloud environments, paving the way for more autonomous and resilient distributed scheduling systems.

Read PDF

Similar papers

Conference Jul 2026

Intelligent Data Engineering Pipelines for Enterprise Applications: Architecture Challenges and Optimization Strategies

The rapid growth of enterprise applications has led to a substantial increase in the volume, velocity, and variety of data, necessitating intelligent data engineering pipelines for efficient processing and analytics. These pipelines play a critical role in enabling scalable data integration, transformation, and delivery across distributed environments. However, existing pipeline architectures face significant challenges, including limited scalability, high latency, inefficient resource utilization, and lack of adaptability to dynamic workloads. Traditional approaches rely on static scheduling and rigid execution models, which restrict their effectiveness in real-time and large-scale scenarios. To address these limitations, this paper proposes an Intelligent Data Engineering Pipeline Architecture supported by an Adaptive Optimization Strategy. The framework integrates hybrid processing, metadata-driven orchestration, and a learning-based cost optimization model to enhance system performance. A key contribution is the Intelligent Adaptive Pipeline Optimization Strategy (IAPOS), which enables dynamic scheduling, cost-aware decision-making, and feedback-driven learning for continuous performance improvement. Experimental evaluation demonstrates that the proposed approach achieves improved scalability, reduced processing latency, and more efficient resource utilization compared to conventional methods. These results highlight the effectiveness of intelligent and adaptive pipeline architectures for next-generation enterprise applications.

D. Bansal, Dinesh Kumar Garg · 0 citations
Review Open access 2025

Intelligent Data Synchronization Techniques for Hybrid Data Platforms

The rapid growth of enterprise systems and cloud computing has transformed data management across hybrid environments integrating on-premise databases, private clouds, and public cloud infrastructures. However, challenges such as data consistency, latency, conflict resolution, security, and fault tolerance remain critical in distributed heterogeneous systems. Traditional synchronization methods are often inadequate for dynamic real-time workloads. This study reviews intelligent data synchronization techniques for hybrid data platforms, emphasizing AI- and machine learning-based approaches that enhance synchronization efficiency, scalability, and reliability. The proposed framework includes four layers: Data Acquisition, Intelligent Synchronization Engine, Adaptive Conflict Management, and Distributed Analytics. Predictive learning algorithms optimize synchronization timing and resource allocation, while adaptive conflict resolution mechanisms minimize inconsistencies. Experimental results show that intelligent synchronization methods reduce delay, improve throughput, enhance scalability, and strengthen failure recovery compared to traditional approaches. The study concludes that AI-driven synchronization is essential for real-time analytics, distributed transactions, and scalable cloud-native applications in modern enterprise environments.

J. Arsac, Gérard Huet · 0 citations
Open access 2024

AI-Assisted Data Pipeline Orchestration for Scalable Analytics

Modern enterprises face increasing demands for scalable and efficient data processing due to rapid data growth. Traditional data pipeline orchestration methods, which rely on static configurations and manual intervention, often lead to inefficiencies in resource use, latency, and fault tolerance. This paper proposes an AI-assisted orchestration framework that integrates machine learning techniques to enable dynamic scheduling, workload prediction, anomaly detection, and resource optimization. By leveraging reinforcement learning, supervised learning, and heuristic methods, the system adapts pipeline configurations in real time based on changing workloads and system conditions. The proposed architecture includes data ingestion modules, AI-driven orchestration engines, adaptive schedulers, and monitoring systems. A key contribution is an intelligent scheduling mechanism that improves execution efficiency and resource utilization. Experimental results show significant improvements over traditional systems, with up to 35% increase in processing efficiency and 25% reduction in latency. The study concludes that AI-driven orchestration is a promising approach for building scalable and autonomous data processing systems, with future work focusing on deeper integration of advanced learning models and real-time adaptability.

J. Weizenbaum, S. Papert · 1 citation
Open access 2025

Self-Adaptive Distributed Computing Models for High-Performance Analytics

This work proposes a scalable, intelligent, and resilient foundation for next-generation high-performance analytics and data-intensive applications that integrates adaptive resource management, intelligent workload scheduling, dynamic task migration, predictive analytics, and machine learning-based optimization to improve computational efficiency and responsiveness.

John Peterson, L. Martínez · 0 citations
Open access 2018

AI-Based Resource Scheduling in Distributed Data Engineering Platforms

Experimental results demonstrate improved resource utilization, reduced scheduling latency, enhanced scalability, lower operational costs, and better workload balancing compared to conventional scheduling approaches, making the framework well-suited for cloud-native and large-scale distributed data engineering environments.

Louis Pouzin · 0 citations
Open access Jul 2026

SPES: A Stochastic Predictive Energy-Aware Scheduling Approach for Efficient Multi-Region Cloud Computing

Cloud computing has transformed the delivery of modern applications and services by providing scalable, flexible, and cost-effective access to computing resources. One of the most critical challenges in cloud environments is the efficient distribution of dynamic workloads across heterogeneous resources, commonly addressed through load balancing and task scheduling techniques. Efficient scheduling plays a vital role in maximizing resource utilization, minimizing response time, and maintaining acceptable Quality of Service (QoS), particularly under dynamic and large-scale workloads. Despite the progress achieved by traditional heuristics such as Min-Min and metaheuristic approaches like the Improved Sparrow Search Algorithm (ISSA), challenges related to scalability, adaptability, and computational overhead remain. Metaheuristic-based approaches often involve iterative optimization processes that may limit their efficiency in real-time scheduling scenarios. In this paper, we propose a lightweight Stochastic Predictive Energy-Aware Scheduling (SPES) algorithm that integrates predictive execution estimation, multi-resource awareness, and stochastic decision-making. Unlike deterministic scheduling strategies, SPES employs a Top K candidate selection mechanism combined with probabilistic weighting and epsilon-greedy exploration to enhance adaptability and avoid suboptimal resource allocation. The proposed method considers CPU, memory, and I/O demands to achieve balanced utilization across heterogeneous hosts while implicitly addressing energy efficiency through utilization-based modeling. The proposed algorithm is implemented and evaluated using the CloudSim 5.0 simulation framework under heterogeneous multi-region cloud environments with varying workload sizes. Experimental results demonstrate that SPES consistently outperforms ISSA and achieves makespan reductions of up to 23.8% while improving scalability, resource utilization, and scheduling efficiency under dynamic cloud workloads. These results indicate that SPES provides an effective lightweight scheduling solution for large-scale and energy-aware cloud computing environments and supports green computing objectives through improved resource efficiency.

M. Yacoub, Ahmed E. Abdel Raouf, Walaa K. Gad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.