This research demonstrates that reliable industrial ML is achieved not by increasing model complexity, but through systematic data-centric practices including structured data preparation, quality-aware pipelines, robustness testing, and adaptive learning, applied across two complementary industrial domains.
Abstract
The use of Machine Learning (ML) is rapidly expanding across diverse scientific and engineering domains. ML offers a powerful advantage over traditional modeling approaches for predictive modeling and analysis of variables of interest. This makes it particularly useful for developing advanced methods in analytical chemistry and residential energy systems, including forecasting (such as predicting hot water demand or chromatographic peak behavior), data quality assessment (such as detecting sensor drift or anomalous consumption patterns), and fault detection (such as identifying heat pump malfunctions or degraded separation performance). While traditional modeling approaches struggle to fully exploit complex, high-dimensional features, existing ML studies in the target domains often rely on limited datasets and lack automated and adaptive frameworks capable of handling the scale, variability, and non-stationary data generated in real operational settings. The availability of large amounts of data in this digital era offers unprecedented opportunities for analysis. However, the successful application of ML depends critically on the quality, preparation, and robustness of ML models and their underlying data. In industrial systems, these requirements are shaped by several interacting factors, particularly feature representation, model robustness, and adaptation under changing conditions. The challenges addressed in this research include data quality management, model selection, robustness assessment, and adaptation under changing real-world conditions. The main contributions of this research are: (i) a semi-automatic data preparation workflow with domain-specific feature engineering for large-scale oligonucleotide chromatography datasets; (ii) an unsupervised quality-centric evaluation framework that automatically clusters input data by quality level without requiring labeled annotations; (iii) the FIUL-Data fault injection framework, which quantifies the resilience boundaries of ML models under controlled data degradation; and (iv) a composite adaptive framework that integrates predictive ML with anomaly detection to enable demand-driven heat pump management in residential energy systems. Together, these contributions demonstrate that reliable industrial ML is achieved not by increasing model complexity, but through systematic data-centric practices including structured data preparation, quality-aware pipelines, robustness testing, and adaptive learning, applied across two complementary industrial domains.
This article delves deep into the confluence of simulation, ML, and statistics, showcasing how they synergize to improve engineering workflows and emphasizes that DCE is not just a technological advancement but a foundational strategy for next-generation engineering solutions.
Benjamin Scott· International Journal of Dat...· 0 citations
A structured methodology for ML-based PdM frameworks, covering data-driven, physics-based, and hybrid approaches, including supervised, unsupervised, and deep learning models is proposed, offering valuable insights for developing efficient and scalable PdM solutions.
Sithik Shah· International Journal of App...· 0 citations
As technology advances, the volume, variety, and velocity of data generation continue to grow, leading to the emergence of big data analytics, which aims to extract valuable insights from these extensive datasets. The increasing volume of data presents several opportunities and challenges. In the context of Statistical Process Monitoring (SPM), analyzing high‐dimensional data can lead to the “curse of dimensionality”, where the data becomes sparse, making it difficult to detect patterns and anomalies. Modeling complex variable relationships is also challenging. Implementing Multivariate SPM (MSPM) in real‐time settings is difficult due to the need for rapid computation and decision‐making, especially when dealing with large volumes of streaming data. A possible solution is to combine MSPM approaches with Machine Learning methods; however, challenges arise related to model interpretability, feature selection, and integration of results. In this paper, we propose a robust method based on a strategy from cluster analysis. The method is compared to multivariate control charts based on the Hotelling statistic and to Dunn's index. An extensive simulation study showed that the proposed method outperforms its competitors when data streams consist of correlated features.
Unknown authors· Quality and Reliability Engi...· 0 citations
This study develops and evaluates an AI-based analytical system for detecting anomalies in industrial processes. The work reviews major sources of risk in industrial control systems, distinguishes point, contextual, and collective anomalies, and summarizes the principal machine-learning approaches used for industrial anomaly detection. A synthetic dataset modeled on the Secure Water Treatment (SWaT) testbed was created with 10 sensor and actuator variables and 10,000 one-second observations, including 1,000 anomalous samples. After missing-value interpolation, duplicate removal, low-variance filtering, and standardization for consistent analysis and visualization, an Isolation Forest with 200 trees was trained in a novelty-detection configuration using normal operating data. On the held-out test set, the model achieved 88.63% accuracy, 45.86% precision, 75.67% recall, and an F1-score of 57.11%. The results show that Isolation Forest can detect most simulated anomalies, although the relatively low precision indicates a substantial false-alarm burden. Future work should validate the approach on authorized real SWaT or PLC-SCADA data, investigate hybrid temporal models, and incorporate explainable-AI methods to support operator decision-making.
Mehdiyeva Almaz, Ahmedov Elmar, Uzakov Gulom et al.· 2026 International Conferenc...· 0 citations
The Industrial Internet of Things (IIoT) is characterized by the generation of vast amounts of time-series data. Modern IIoT systems enable efficient collection, storage, and querying of massive industrial time-series data, making the processing and analysis of such data a key enabler for data-driven decision-making in modern manufacturing. To provide researchers and practitioners with comprehensive guidance on industrial time series data analysis, this paper presents a systematic review of state-of-the-art methods—spanning statistical approaches, machine learning (ML), deep learning (DL), and cutting-edge large models—along with their applications in industrial decision-making. It details the application status of these methods in key equipment condition monitoring, manufacturing process supervision, and energy network management. Additionally, the paper discusses existing gaps between methods and real-world applications, as well as future trends and challenges, such as optimizing data structures for cost-sensitive learning, exploring causality and time-series-oriented model architectures, and developing cascaded/hybrid pipelines for end-to-end industrial use cases. Ultimately, this review aims to inspire innovations in realizing data-driven intelligent decision-making for next-generation IIoT systems.
Li-Lan Liu, Yixiang Zhang, Yanning Sun et al.· ACM Computing Surveys· 0 citations
Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the need for real-time decision-making across a wide range of timescales. Addressing these challenges traditionally relies on model-based methods that, while accurate, can be too slow for operational demands. Machine learning (ML) has therefore emerged as a faster, data-driven alternative. As grid topology plays a central role in power system operation, graph machine learning (GML) methods offer a natural framework for incorporating topological dependencies as an inductive bias. We survey nearly 800 papers at the intersection of GML and power systems, covering forecasting, state estimation, optimization, control, fault diagnosis, and cybersecurity. Power systems constitute an unusually rich benchmark setting for GML, as they combine hard physical constraints, multi-scale dynamics, safety-critical requirements, and scarce labeled data within a single, well-defined domain. Conversely, power systems can benefit from utilizing GML to complement classical solvers, as GML provide scalable, topology-aware approximations with promising generalization and computational efficiency. We identify open challenges, including limited real-world deployment and the need for interpretable models in safety-critical settings. Despite the rapidly growing number of publications, standardized benchmarks and open datasets remain scarce, leaving many results difficult to reproduce and undermining the long-term scientific credibility of the field. We further derive a structured requirements catalog for ML-ready power grid benchmarks, intended to guide future dataset development and improve reproducibility across studies. We call on the community to prioritize dedicated benchmark studies and the release of open datasets and models.
Martin Sadric, Sebastian Pütz, Christian Nauck et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.