This work addresses the resulting Multi-Objective Flexible Job Shop Scheduling Problem by proposing a deep reinforcement learning framework that jointly minimizes makespan and energy cost, and evaluates the approach against NSGA-II and Joined Heuristics on synthetic instances.
Abstract
Rising energy costs and the increasing share of renewable generation create incentives to align production schedules with dynamic electricity prices and on-site solar generation. We address the resulting Multi-Objective Flexible Job Shop Scheduling Problem by proposing a deep reinforcement learning framework that jointly minimizes makespan and energy cost. A single preference-conditioned policy approximates the Pareto front at inference time, eliminating the need to train separate models for different objective weightings. The agent acts as a hyper-heuristic, selecting among heuristic actions at each decision point, including strategies that intentionally delay operations to exploit periods of lower electricity prices or higher solar generation. Preferences are integrated throughout the network via Feature-wise Linear Modulation, while a dual-critic architecture and a diversity loss preserve preference-specific policy behaviors. We evaluate the approach against NSGA-II and Joined Heuristics on synthetic instances ranging from 10×5×5 to 15×15×15 jobs, operations per job, and machines using normalized hypervolume and inverted generational distance. While NSGA-II performs best on the smallest instances, the proposed approach becomes increasingly competitive as problem size grows and achieves the best results on the largest evaluated instances. These findings indicate promising scalability within the investigated problem range.
PGMPO is proposed, a novel learning framework consisting of a simple but effective multi-policy modeling approach that allows a single network to represent multiple decision-makers, and a preference-driven model optimization method that effectively guides policies to learn diverse and specialized problem-solving strategies without the need for explicit reward functions.
Inguk Choi, Woo-Jin Shin, Sang-Hyun Cho et al.· 0 citations
Modern manufacturing requires scheduling methods that adapt to changing order arrivals, machine disruptions, customer priorities, stakeholder preferences, and time-varying energy conditions. This paper proposes a preference-conditioned deep reinforcement learning (DRL) approach for dynamic scheduling in sustainable and robust manufacturing. The approach is embedded in a cyber-physical production system (CPPS)-oriented framework that links production states, machine availability, energy-related background data, simulation-based learning, performance monitoring, and decision support. Within this framework, a Double Deep Q-Network (DDQN) scheduler is developed for joint job sequencing, machine assignment, and start-time adjustment. The scheduler uses a candidate-based state representation for dynamic order arrivals, vector-valued Q-output for objective-specific value estimation, and a priority- and preference-aware reward design. Customer priorities are treated as order-level attributes, while stakeholder preferences are encoded as system-level objective weightings. This enables one policy to consider energy-related cost, carbon emissions, energy demand, and tardiness while adapting to different preference profiles. The concept is demonstrated in an on-demand manufacturing (ODM)-oriented parallel CNC machining case with heterogeneous orders, product-specific setup and processing requirements, hourly electricity prices, carbon-intensity signals, and curriculum-adaptive machine breakdowns. DDQN is compared with three dispatching rules and two DRL baselines under shared training and testing scenarios. The results show that DDQN achieves the lowest energy-related cost and carbon emissions in training and unseen testing while maintaining acceptable delivery performance. Overall, the study demonstrates the potential of CPPS-oriented and preference-conditioned DRL for adaptive, energy-aware, and robust scheduling in smart manufacturing systems.
Chao Zhang, Gabriela Ventura Silva, Christoph Herrmann· Production Engineering· 0 citations
This paper proposes ALP-PPO, an adaptive Lagrangian penalty-enhanced proximal policy optimization algorithm for real-time rescheduling under concurrent machine breakdowns and rush orders, and indicates that the adaptive Lagrangian mechanism reduces constraint violations by more than 40% relative to fixed-penalty alternatives while keeping the primary objectives competitive.
Analysis of queue-level dynamics reveals more regular behavior in the evaluated scenarios, with reduced fluctuations in queue lengths, batch waiting, and minimum queue entropy over time, indicating that the proposed ABC-based approach can improve observed predictability at the queue level.
The increasing penetration of distributed renewable energy sources has intensified the need for intelligent bidding strategies in virtual power plants (VPPs), where reliable communication and real-time information exchange are essential for coordinated energy management. This study proposes an optimal bidding path construction framework based on the Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm for VPP participation in electricity spot markets. A Markov decision process is established to characterize dynamic market interactions, and customized state-space optimization, constrained action-space design, and a multi-objective reward function are integrated into the Actor–Critic architecture to jointly maximize economic returns while satisfying operational constraints. The framework further incorporates communication-aware resource coordination mechanisms that leverage edge computing and low-latency information exchange to enhance decision consistency under uncertain renewable generation and volatile market conditions. Experimental evaluation demonstrates that the improved DDPG algorithm increases average daily revenue by 39.1% compared with conventional DDPG, accelerates convergence by approximately 15%, reduces revenue volatility by 12%, and maintains the constraint violation rate at 1.2%. In addition to intelligent energy scheduling, the proposed methodology provides valuable insights into communicationenabled power systems, distributed electromagnetic information networks, and wireless coordination infrastructures requiring adaptive decision-making and reliable multi-node information
P. Hu, N. Shen, H. Guo et al.· Advanced Electromagnetics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.