Skip to content
Open access

Uncertainty-Aware Predictive Maintenance Scheduling: A Decision-Support Framework for Industrial Production Systems

Sep 2026 · Journal of Engineering, Project, and Production Management · Vol 16, pp. 1-10 · 0 citations

TL;DR

This paper describes a decision-support framework that connects data-driven prognostics to a production-constrained optimization model and evaluated the framework on a public milling benchmark and eight months of data from a twelve-machine packaging facility.

Abstract

Preventive maintenance in a manufacturing plant is, at bottom, a resource-allocation problem: every intervention weighs the cost of replacing a component too early against the risk of an unplanned stoppage under the challenging realities of production calendars, crew availability, and machine access. Most facilities still rely on fixed calendar intervals that ignore what the equipment is signaling, paying for the difference in wasted component life and avoidable downtime. This paper describes a decision-support framework that connects data-driven prognostics to a production-constrained optimization model. A Gradient-Boosted Regression Tree (GBRT) supplies point estimates of Remaining Useful Life (RUL). A Quantile Regression Forest (QRF) turns these into calibrated prediction intervals with finite-sample coverage guarantees, and a Mixed-Integer Linear Program (MILP) converts the resulting per-machine failure probabilities into maintenance schedules that respect shift windows and crew capacity. The contribution lies less in the individual components than in how they are joined: uncertainty flows directly into the scheduler, so every intervention decision reflects both predictive risk and production feasibility. We evaluated the framework on a public milling benchmark and eight months of data from a twelve-machine packaging facility (ten independent runs each). Relative to the calendar-based policy currently used at the facility, total maintenance cost fell by 6.3% on the packaging line and 11.9% on the NASA milling dataset; unplanned downtime events fell by 7.2% and 44.4%; and unnecessary preventive replacements fell by 26.2% and 66.7%. Calibration error remained below 1.5 percentage points at every coverage level tested, and all scheduling instances were solved to certified optimality within seconds for fleets of up to 50 machines on standard hardware.

Read PDF

Similar papers

Open access Aug 2026

Optimising Aircraft Fleet Maintenance with Reinforcement Learning: A numerical case study

Unexpected failures and inefficient maintenance plans lead to costly downtime in many industrial fields, undermining operational sustainability and overall profitability, or, in the worst scenarios, creating unsafe situations for people. These issues justify a shift towards predictive maintenance, and structural health monitoring (SHM) has emerged as a viable option. By exploiting modern hardware and software technologies, the diagnosis and prognosis of monitored structures can be performed while accounting for estimation uncertainty. In addition, for complex structures, monitoring all components is not feasible for technical or economic reasons, and a mix of condition and schedule-based maintenance approaches is still required. To fully exploit the SHM benefits, an effective decision-making process is necessary, including Opportunistic Maintenance (OM) that accounts for both scheduled stops and the Residual Useful Life (RUL) predicted by prognostic models. In this context, decision-making becomes significantly more complex when managing an entire fleet rather than a single asset, reflecting the real-world challenges companies face. This study explores the application of Reinforcement Learning (RL) to automated decision-making in aircraft fleet maintenance, aiming to minimise operational and maintenance (O&M) life-cycle costs. The RL agent is trained and deployed within a fleet life cycle simulator that replicates real-world conditions, including mission schedules, maintenance operations, repair bases with varying capabilities, and potential unforeseen events. Each aircraft is modelled as an ensemble of subcomponents, with only a portion equipped with an SHM system that provides Remaining Useful Life (RUL) predictions, expressed as probability distributions, and fault detection. The agent leverages this data to develop an optimal maintenance policy, aligning scheduled stops for non-monitored components with necessary interventions for monitored parts. Specifically, it determines maintenance schedules and selects which components to replace based on fleet-wide data, expected mission schedules, and aircraft availability. While the study is demonstrated using a fleet life-cycle simulator, it serves as a preliminary step toward the real-world implementation of autonomous agents. The integration of the agent within the fleet simulator, potentially based on a real scenario, enhances its capabilities as a Digital Twin, supporting management decisions. Furthermore, it leverages the potential benefits of SHM technology, encouraging broader industry adoption.

Emanuele Petriconi, Claudio Giglio, marco, sbarufatti · 0 citations
Open access Aug 2026

Integrated Decision-Making Architecture Merging Predictive Maintenance, Risk Assessment, and Revenue Optimization in Rental Fleet Management

Component degradation that stays below the OBD-II diagnostic trouble code threshold escapes fixed-threshold maintenance scheduling in rental fleets. Revenue management literature classifies vehicle condition as an exogenous constraint on capacity allocation. The economic cost of maintenance timing sits outside the optimization objective in most published formulations. This paper specifies an architecture in which component-health estimation, behavioral risk scoring, and revenue-lifecycle optimization share a single time-indexed vehicle state. Maintenance timing resolves through a demand-aware window search subordinate to a hard safety override; an explicit priority hierarchy addresses the failure mode a purely multiplicative composite score produces when a single factor approaches zero. Evaluated through an 18-month field pilot across 1200 vehicles in the Phoenix, Arizona metropolitan market, the architecture achieves 91.4% empirical interval coverage on calibrated remaining-useful- life estimates, close to the 90% nominal target. Component degradation surfaces a median of 11.3 days before functional failure on cases that stay below a diagnostic trouble code threshold. Behavioral risk estimates converge within 2.6 minutes when cross-session history exists; a kinematics- only baseline converges at 14.1 minutes. Opportunity cost per maintenance event drops by 29.5% under demand coefficients of variation at or above 0.35, a benefit that narrows to 6.8% once demand approaches uniformity. Prior insurance-telematics findings require multi-month accumulation for stable risk scoring. An open question remains: whether cumulative behavioral risk exposure should enter the lifecycle score as a temporal pattern, distinguishing accumulation trajectories that reach the same aggregate value through different event sequences.

V. Kolesnykov · 0 citations
Open access Jul 2026

An integrated framework for production lot sizing and predictive maintenance based on Remaining Maintenance Life and Proportional Hazard Models

Equipment reliability is critical to maintaining production efficiency and controlling manufacturing costs. Predictive maintenance (PdM), based on machine condition prediction, can effectively reduce the risk of machine failure; consequently, machine degradation assessment and the prediction of remaining maintenance life (RML) are crucial for maintenance decision-making. Moreover, because equipment condition affects production planning, PdM should be integrated into the traditional economic production quantity (EPQ) model. The main contribution of this study is the introduction of RML as a decision-support metric linking equipment degradation prediction with production lot-sizing decisions, thereby enabling the joint optimization of EPQ and PdM policies. To minimize expected average cost, maintenance decisions are integrated into the production lot-sizing model to determine the optimal production lot size and maintenance policy. This study considers a single-machine production process in which the ARMA method is used to forecast the machine degradation index. Cox’s proportional hazard model (PHM) is then employed to estimate machine reliability based on the predicted degradation index. Based on this reliability assessment, remaining maintenance life (RML), rather than the traditional remaining useful life (RUL), is employed to link degradation prediction with the EPQ model and to characterize the machine deterioration process. An integrated EPQ–PdM model is developed to jointly determine the optimal production lot size and maintenance policy while minimizing the expected average cost (EAC) over the production cycle. Finally, a case study of an automotive bumper factory demonstrates the effectiveness of the proposed framework. The results show that the framework reduces EAC by 19.6 % relative to the current production strategy and identifies an optimal maintenance threshold of Rsafe = 0.4 with six maintenance cycles.

Ci-Wen Zhong, L. Lei, H. Zhang et al. · 0 citations
Preprint Aug 2026

Integrating Prognostics, Maintenance, and Tail Assignment under Remaining Useful Life Uncertainty: A Stochastic Optimisation Approach for Airline Reliability

Ensuring reliability, safety, and economic efficiency in airline operations requires maintenance and fleet scheduling strategies that explicitly account for uncertainty in Remaining Useful Life (RUL) predictions. However, the integration of prognostic uncertainty into operational decision-making remains a major challenge. In practice, tail assignment (TA) and maintenance scheduling (MS) are typically optimized separately or sequentially, thereby limiting the effective use of predictive health information despite their strong interdependencies. This paper proposes a unified optimisation framework that jointly integrates TA, MS, and predictive maintenance (PdM) under RUL with confidence intervals. The problem is formulated as a stochastic mixed-integer linear program, and a scalable solution approach is developed by embedding a neural network surrogate to approximate expected disruption costs resulting from RUL uncertainty. The proposed framework is evaluated using operational scenarios derived from real-world airline data. Results show that explicitly incorporating prognostic uncertainty in a joint planning model reduces operational risk, i.e., downstream disruption costs and flight cancellations, compared to deterministic and sequential approaches, at the expense of moderate increases in planning cost. These findings highlight the value of tightly coupling predictive maintenance with operational planning and demonstrate the potential of surrogate-assisted stochastic optimisation for scalable, uncertainty-aware airline decision-making.

Benno Käslin, Marta Ribeiro, D. Zarouchas et al. · 0 citations
Jul 2026

A hybrid framework for data-driven predictive maintenance: probabilistic RUL estimation and early failure signaling via control charts

The purpose of this study is to develop and validate a robust framework for data-driven predictive maintenance (PdM) that estimates the remaining useful life (RUL) of equipment and signals the optimal time to initiate maintenance activities. By integrating statistical modeling and machine learning techniques, the proposed framework aims to minimize unplanned downtimes, reduce maintenance costs and enhance operational efficiency. It addresses critical gaps in existing methods by providing real-time condition monitoring and predictive alerts, enabling maintenance personnel to make informed decisions and optimize maintenance scheduling. This study introduces a data-driven predictive maintenance framework designed to estimate the RUL of equipment and provide timely signals for initiating maintenance tasks. The framework utilizes a combination of statistical modeling and machine learning algorithms. A Weibull distribution with time-varying parameters is employed to model the RUL, while random forest and exponentially weighted moving average (EWMA) control charts are integrated for failure prediction and maintenance signaling. The methodology is validated using a synthetic dataset that simulates real-world scenarios, enabling robust evaluation of the proposed approach in terms of accuracy and reliability. The findings demonstrate that the proposed predictive maintenance framework effectively estimates the RUL of equipment with high accuracy and provides timely signals for initiating maintenance tasks. Validation on a synthetic dataset reveals that the framework consistently predicts failure probabilities and generates alerts well in advance, allowing sufficient time for maintenance planning. The integration of Weibull distribution modeling and Random Forest classifiers enhances the reliability of predictions, while the use of EWMA control charts ensures robust monitoring of failure probabilities. These results highlight the framework's potential to reduce downtimes and maintenance costs. This study introduces a novel data-driven predictive maintenance framework for accurately estimating the RUL of equipment and providing timely signals for maintenance actions. Unlike existing methods, the proposed approach integrates a probabilistic model with machine learning techniques, leveraging time-varying Weibull distributions and advanced statistical tools for robust predictions. The framework not only improves failure prediction accuracy but also enhances maintenance planning by signaling appropriate times for action. This contributes to reducing downtime and costs while increasing operational efficiency, offering a significant advancement for Industry 5.0 and smart manufacturing applications.

Ahmad Razavi, M. R. Rasouli, M. Pishvaee et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.