Skip to content

Reinforcement learning-based decision making for sustainable manufacturing operations

2026 · Materials Research Proceedings · Vol 71, pp. 81-88 · 0 citations

TL;DR

Comparative analysis shows that the RL-based approach outperforms the rule-based and heuristic strategies and reports remarkable energy efficiency and operational sustainability.

Abstract

Abstract. Sustainable manufacturing involves being able to optimize productivity, energy efficiency, and environmental impact simultaneously given dynamic and uncertain operating conditions. The traditional optimization methods are unadaptable and cannot easily reflect the real-time changes in the system. This paper provides a sophisticated reinforcement learning (RL)-based decision-making model of sustainable manufacturing process. The manufacturing system is modelled as a Markov Decision Process (MDP) and a Deep Q-Network (DQN) is used to learn about the optimal control policies by interacting with the environment continuously. Multi-objective reward function is created to include production rate, energy usage, machine usage and minimization of waste. The suggested framework is tested in a virtualized smart factory setting, where the demand is stochastic and machines have variability. Comparative analysis shows that the RL-based approach outperforms the rule-based and heuristic strategies and reports remarkable energy efficiency and operational sustainability. The findings prove RL as a potential solution to adaptive and intelligent manufacturing control.

View source

Similar papers

2026

Reinforcement learning-based adaptive control strategies for sustainable production systems

Abstract. The need to achieve sustainable production has become an urgent necessity in the conditions of stricter environmental requirements and the rise in the cost of energy worldwide. The classical proportionalintegralderivative controllers and linear Model Predictive Controllers are conventional model-based control strategies that by nature rely on precise process models and fixed optimization horizons and hence are not well suited to the dynamic, non-linear, and complex nature of the modern production environment. The current paper suggests a new adaptive control system based on Reinforcement Learning (RL), where the overall production system is optimized and controlled to achieve Sustainable Production Systems, with the multi-objective rewarding function, which aims to minimize the Overall Equipment Effectiveness (OEE), defect rate, specific energy consumption, and carbon dioxide emissions, and a physics-informed Digital Twin safety filter that stops unsafe policy execution in training and deployment. The proposed framework is evaluated on a multi-machine flexible manufacturing cell benchmark, which includes CNC milling, turning, and robotic assembly, and yields an OEE of 93.6, a defect rate of 0.9, a decrease in specific energy consumption of 27.8, and a decrease in CO 2 emissions of 23.4 compared to PID baseline controllers. Experiments with ablation prove all the above-mentioned in the necessity of each of the architectural constituents, and the Pareto frontier analysis proves that the proposed SAC agent is the best in the OEE-versus-energy trade-off space compared to all of the competing methods.

Apoorva Verma · 0 citations

Deep reinforcement learning-based optimal energy management for microgrids

Microgrids play a critical role in enhancing the flexibility, reliability, and sustainability of modern power systems by integrating distributed energy resources, energy storage systems, and controllable loads. However, the inherent uncertainty of renewable generation and the stochastic nature of load demand pose significant challenges to optimal energy management. To address these issues, this paper proposes a deep reinforcement learning (DRL)-based optimal energy management framework for microgrids. The problem is formulated as a Markov decision process, where the system state captures renewable generation, load demand, and storage status, while the control actions determine power dispatch and energy storage operation. A deep reinforcement learning model is developed to learn optimal control policies through continuous interaction with the environment, enabling adaptive decision-making under dynamic and uncertain conditions. To improve learning efficiency and policy stability, state normalization and reward shaping strategies are incorporated. Furthermore, a constrained optimization mechanism is introduced to ensure operational safety and economic feasibility. Experimental results on benchmark microgrid scenarios demonstrate that the proposed method outperforms conventional rule-based strategies and model-based optimization approaches in terms of operational cost reduction, energy utilization efficiency, and robustness under uncertainty. The results indicate that the proposed DRL-based framework provides an effective and scalable solution for intelligent microgrid energy management.

Li Chen, Hong-Qiao Li, Zhen-Xing Chen et al. · 0 citations
Open access Aug 2026

Preference-conditioned deep reinforcement learning for dynamic scheduling in sustainable and robust manufacturing

Modern manufacturing requires scheduling methods that adapt to changing order arrivals, machine disruptions, customer priorities, stakeholder preferences, and time-varying energy conditions. This paper proposes a preference-conditioned deep reinforcement learning (DRL) approach for dynamic scheduling in sustainable and robust manufacturing. The approach is embedded in a cyber-physical production system (CPPS)-oriented framework that links production states, machine availability, energy-related background data, simulation-based learning, performance monitoring, and decision support. Within this framework, a Double Deep Q-Network (DDQN) scheduler is developed for joint job sequencing, machine assignment, and start-time adjustment. The scheduler uses a candidate-based state representation for dynamic order arrivals, vector-valued Q-output for objective-specific value estimation, and a priority- and preference-aware reward design. Customer priorities are treated as order-level attributes, while stakeholder preferences are encoded as system-level objective weightings. This enables one policy to consider energy-related cost, carbon emissions, energy demand, and tardiness while adapting to different preference profiles. The concept is demonstrated in an on-demand manufacturing (ODM)-oriented parallel CNC machining case with heterogeneous orders, product-specific setup and processing requirements, hourly electricity prices, carbon-intensity signals, and curriculum-adaptive machine breakdowns. DDQN is compared with three dispatching rules and two DRL baselines under shared training and testing scenarios. The results show that DDQN achieves the lowest energy-related cost and carbon emissions in training and unseen testing while maintaining acceptable delivery performance. Overall, the study demonstrates the potential of CPPS-oriented and preference-conditioned DRL for adaptive, energy-aware, and robust scheduling in smart manufacturing systems.

Chao Zhang, Gabriela Ventura Silva, Christoph Herrmann · 0 citations
2026

Machine learning–based predictive control for energy-efficient manufacturing systems

Abstract. The fact that operational costs and the environmental impact are increasing is what has made energy consumption in manufacturing systems a serious issue. In this paper, a machine learning (ML)-based predictive control model is introduced to enhance energy efficiency in the contemporary manufacturing settings. The offered solution combines predictive models based on data and Model Predictive Control (MPC) to optimize the performance of the systems in real time. Machine learning algorithms are used to predict the energy demand, process dynamics, and disturbances, as well as to make decisions proactively. Industrial case studies confirm the validity of the framework, showing great progress in terms of energy efficiency, productivity, and stability of operations.

Prabhakara Rao Kapula · 0 citations
Conference Aug 2026

AI-Driven Predictive Supply Chain Optimization Using Reinforcement Learning Algorithms

In a dynamic market which is getting more and more uncertain, efficient supply chain management has emerged as a challenge of utmost importance to organizations. Conventional optimization methods are not usually capable of adapting to the changing demands in real-time and complicated operational constraints. The paper introduces a predictive supply chain optimization model that uses AI and is developed on the basis of a new Hybrid Deep Reinforcement Learning algorithm (HDRL-SCO). The proposed solution will combine the application of Long Short-Term Memory (LSTM) networks to make precise demand forecasts with a Deep Q-Network (DQN) to make smart decisions. The system is able to learn the best policies to manage inventory, transportation planning, and order fulfillment, by constantly interacting with the environment by modelling the supply chain as a Markov Decision Process. As the experimental outcomes show, the suggested model will result in a considerable decrease in the overall operating costs, a higher level of services, increased inventory turnover as well as a decrease in the delays in the fulfillment process as compared to the traditional approaches. The framework is highly adaptive, scalable, and resilient to uncertain and dynamic environments and thus can be applied in the real-life supply chain.

B. Rajnarayanan, C. R. Usha, Venkata Appaji Sirangi et al. · 0 citations
2026

Intelligent energy management of manufacturing systems integrated with renewable sources

Abstract. The growing presence of renewable energy sources into manufacturing systems presents massive challenges in their intermittency and uncertainty, causing inefficiency in energy consumption and production planning. The paper introduces a smart energy management network of manufacturing systems that dynamically balances machine activities with the availability of renewable energy. The suggested solution consists of a predictive model of renewable generation and an optimization-based scheduling system to reduce the cost of energy, delays in production, and grid reliance. There is a multi-objective formulation that is created based on energy use, operating limitations and emission parameters. A smart decision-making policy, founded on machine learning-aided prediction and adaptive scheduling, is applied to guarantee real-time responsiveness when operating in different energy conditions. The framework is tested on a realistic load and renewable profile on a representative manufacturing scenario. Findings show that there are considerable increases in the use of renewable energy, decrease in peak grid demand, and general energy cost savings over traditional scheduling methods. The new methodology provides a scalable and viable approach to sustainable and energy-efficient smart manufacturing systems.

A. R · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.