Skip to content
Open access

Refined Environmental Design for Reinforcement Learning Framework in Job Shop Scheduling

2026 · IEEE Access · Vol 14, pp. 110314-110330 · 0 citations · 62 references
Computer Science

Abstract

Machine failures, which represent forms of performance degradation, are common in real-world manufacturing systems, however, they are often overlooked in job shop scheduling solutions that primarily focus on complete machine breakdowns. These subtle disruptions can lead to cascading delays and reduced system efficiency. This study proposes a reinforcement learning (RL) framework designed to address the Job Shop Scheduling Problem (JSSP) in environments affected by such failures. Unlike traditional RL-based scheduling models that concentrate on total breakdowns, this work considers more nuanced disruptions, such as processing slowdowns, which frequently occur in practical settings. The proposed framework enhances a Q-learning algorithm by introducing a refined environment that incorporates an extended state representation and a reward function tailored to account for performance degradation. These enhancements enable the RL agent to learn adaptive scheduling policies that minimize makespan while effectively responding to partial machine failures. The framework is validated using the Taillard benchmark dataset across varying levels of disruption and job-machine configurations. Experimental results show that the proposed environment consistently delivers superior scheduling performance compared to baseline models that do not consider machine failures. The findings highlight the framework’s potential to improve scheduling resilience and efficiency in dynamic production environments.

Read PDF

Similar papers

Preference-Guided Multi-Policy Optimization for Flexible Job Shop Scheduling

PGMPO is proposed, a novel learning framework consisting of a simple but effective multi-policy modeling approach that allows a single network to represent multiple decision-makers, and a preference-driven model optimization method that effectively guides policies to learn diverse and specialized problem-solving strategies without the need for explicit reward functions.

Inguk Choi, Woo-Jin Shin, Sang-Hyun Cho et al. · 0 citations
Book Open access Aug 2026

EDA Job Scheduling Using Reinforcement Learning with Adaptive Macro Actions

Job scheduling in Electronic Design Automation (EDA) environments presents unique challenges due to high-frequency job submissions, short job durations, and strict latency requirements. Production schedulers such as IBM Spectrum LSF employ robust heuristics like First-Come-First-Served (FCFS) that provide predictable behavior and fairness guarantees. However, these rule-based approaches do not learn from historical workload patterns, leaving potential for further optimization through adaptive methods. We present Adaptive Job Selection (AJS), a reinforcement learning-based agent that learns to schedule pending jobs for reduced job waiting and completion time on LSF clusters for EDA workloads. AJS introduces macro actions that dispatch multiple jobs per inference, addressing the credit assignment problem inherent in high-frequency scheduling while meeting real-time latency constraints. Our lightweight neural network architecture employs cross-attention to capture interactions between job buckets and cluster state, enabling inference at the frequency required by EDA workloads. We deploy AJS as an external plugin in IBM Spectrum LSF and evaluate it on an IBM LSF cluster. Experiments demonstrate that AJS achieves a 62.6% reduction in average job waiting time, a 19.8% reduction in job completion time, a 16.0% improvement in job throughput, and 3.05 percentage points higher CPU utilization compared to the default scheduler. We also share practical lessons for bridging the simulation-to-production gap. To our knowledge, AJS is the first open-source, deployable RL-based scheduler designed for production EDA environments.

Yiming Shao, Aijun An, Michael Spriggs et al. · 0 citations
Open access Aug 2026

Preference-conditioned deep reinforcement learning for dynamic scheduling in sustainable and robust manufacturing

Modern manufacturing requires scheduling methods that adapt to changing order arrivals, machine disruptions, customer priorities, stakeholder preferences, and time-varying energy conditions. This paper proposes a preference-conditioned deep reinforcement learning (DRL) approach for dynamic scheduling in sustainable and robust manufacturing. The approach is embedded in a cyber-physical production system (CPPS)-oriented framework that links production states, machine availability, energy-related background data, simulation-based learning, performance monitoring, and decision support. Within this framework, a Double Deep Q-Network (DDQN) scheduler is developed for joint job sequencing, machine assignment, and start-time adjustment. The scheduler uses a candidate-based state representation for dynamic order arrivals, vector-valued Q-output for objective-specific value estimation, and a priority- and preference-aware reward design. Customer priorities are treated as order-level attributes, while stakeholder preferences are encoded as system-level objective weightings. This enables one policy to consider energy-related cost, carbon emissions, energy demand, and tardiness while adapting to different preference profiles. The concept is demonstrated in an on-demand manufacturing (ODM)-oriented parallel CNC machining case with heterogeneous orders, product-specific setup and processing requirements, hourly electricity prices, carbon-intensity signals, and curriculum-adaptive machine breakdowns. DDQN is compared with three dispatching rules and two DRL baselines under shared training and testing scenarios. The results show that DDQN achieves the lowest energy-related cost and carbon emissions in training and unseen testing while maintaining acceptable delivery performance. Overall, the study demonstrates the potential of CPPS-oriented and preference-conditioned DRL for adaptive, energy-aware, and robust scheduling in smart manufacturing systems.

Chao Zhang, Gabriela Ventura Silva, Christoph Herrmann · 0 citations
Open access 2026

Entropy-Regulated Job-Shop Scheduling: A Bottom–Up Artificial Bee Colony Algorithm for Semiconductor Manufacturing

Analysis of queue-level dynamics reveals more regular behavior in the evaluated scenarios, with reduced fluctuations in queue lengths, batch waiting, and minimum queue entropy over time, indicating that the proposed ABC-based approach can improve observed predictability at the queue level.

Elnaz Khatmi, Khalil Al-Rahman Youssefi, Wilfried Elmenreich · 0 citations
Open access Jul 2026

Preference-Conditioned Reinforcement Learning for Energy-Aware Multi-Objective Flexible Job Shop Scheduling

This work addresses the resulting Multi-Objective Flexible Job Shop Scheduling Problem by proposing a deep reinforcement learning framework that jointly minimizes makespan and energy cost, and evaluates the approach against NSGA-II and Joined Heuristics on synthetic instances.

Dustin Moreira Simoes, Marvin Brune, Mehmet Ulrich et al. · 0 citations
Open access Aug 2026

Adaptive Lagrangian Penalty-Enhanced Proximal Policy Optimization for Flexible Job Shop Rescheduling with Worker Workload Constraints Under Concurrent Dynamic Disturbances

This paper proposes ALP-PPO, an adaptive Lagrangian penalty-enhanced proximal policy optimization algorithm for real-time rescheduling under concurrent machine breakdowns and rush orders, and indicates that the adaptive Lagrangian mechanism reduces constraint violations by more than 40% relative to fixed-penalty alternatives while keeping the primary objectives competitive.

Yuanmeng Zhou, Haoyi Tan, Jiawei Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.