Automated Machine Learning for IoT Intrusion Detection: A Comparative Evaluation of FLAML and TPOT Under Multiple Validation Strategies
Susan Al Naqshbandi
Aug 2026· Sulaimani Journal for Engineering Sciences· Vol 12, pp. 205-211· 0 citations
TL;DR
A fully automated IoT-based Network Intrusion Detection System (NIDS), utilizing the Gotham Dataset 2025, using two AutoML approaches: TPOT (Tree-based Pipeline Optimization Tool) and FLAML (cost-aware lightweight AutoML framework).
Abstract
The vast expansion of IoT devices significantly expands the potential attack surface for networks; therefore, the increased complexity in detecting intrusions from these new sources is caused by the large dimensionality of the data, the severe class imbalance issue that exists when defining a normal versus abnormal behavior model based upon this data and the chaotic nature of network traffic.
This paper proposes a fully automated IoT-based Network Intrusion Detection System (NIDS), utilizing the Gotham Dataset 2025. Two AutoML approaches are used as part of the proposed system: TPOT (Tree-based Pipeline Optimization Tool) and FLAML (cost-aware lightweight AutoML framework). Both systems were evaluated using four different validation methods against approximately 5.8 million network traffic instances. FLAML achieved better results than TPOT in all evaluations, including achieving 99.98% accuracy at one evaluation. TPOT achieved comparable or slightly better performance than FLAML for precision and recall, but was less stable in its performance when dealing with class imbalance.
The majority of assaults in heterogeneous networks are detected by intrusion detection systems (IDS). Cyberattack kinds that seriously harm networks are difficult for conventional IDSs to detect. The majority of existing solutions rely on deep learning models, which have a significant computational and energy overhead that limits their use in IoT environments with limited resources. A lightweight IDS based on ML is proposed in this research as a solution to this difficulty. Predicting the behavior of network traffic is achieved using ToN-IoT data and a tailored preprocessing pipeline. The voting-based ensemble classifier is built through the combination of models of RF and LightGBM to enhance the stability of the classification. The standard performance measures that are utilized to evaluate the proposed approach include accuracy, precision, recall, F1score, false alarm rates, and ROC analysis. The experimental findings indicate that RF achieve 99.81% accuracy, LGBM achieve 99.83%, and the ensemble model has a high accuracy of 99.99% with very low false alarms. Comparative evaluation with traditional ML and DL models demonstrates improved detection reliability with reduced computational overhead. These results prove that the suggested architecture is both computationally efficient and practically applicable to IoT settings with limited resources. However, direct hardware-level energy measurements are required to fully quantify the energy-saving characteristics of the proposed IDS.
Abhinay Kumar Reddy Seella, Rupesh Shirke, Vijay Kumar Kasuba et al.· International Conference on...· 0 citations
This study investigates the effectiveness of supervised machine learning techniques for detecting cyberattacks in IoT-based smart city networks using the TON_IoT dataset, finding that advanced ensemble learning combined with robust feature engineering provides a reliable and scalable solution for securing smart city IoT networks.
E. Okonta, Oluwaseun Bamgbose· ABC2: Journal of Architectur...· 0 citations
The rapid propagation of Internet of Things (IoT) devices has significantly expanded the cyber-attack surface, particularly in essential infrastructure sectors such as energy, water, and healthcare. Machine learning (ML) based intrusion detection systems (IDS) offer a promising defense, but their real-world deployment is often hindered by data imbalance, lack of interpretability, and computational demands. In this paper, we introduce a lightweight ensemble approach, which integrates XGBoost and LightGBM using a soft-voting method. The system is evaluated on the IDSAI dataset after eliminating duplicates, resulting in 693,116 unique samples with a natural class imbalance. The preprocessing phase includes data cleansing and data scaling. The results indicate that the proposed ensemble achieves 99.95% accuracy, 99.95% F1-score, and a perfect AUC of 1.0 on a test set of 207,935 samples. Training completes in under 8 seconds on a standard CPU. The feature importance (gain) highlights delta_time; packet inter-arrival time, as the most significant feature, followed by source/destination ports. SHapley Additive exPlanations (SHAP) analysis provides local explanations, revealing that high inter-arrival times push predictions toward malicious—likely due to slow scanning or burst-and-pause attack patterns. All code and the trained model are publicly available to facilitate reproducibility1.
Nooruddine F. Assarwie, F. Alqasemi, Tasnim M. Al-Khawlani et al.· 2026 6th International Confe...· 0 citations
The explosion of Internet of Things (IoT) deployment over the past decade has served as a foundational pillar for global digital transformation. However, the rapid expanding attack surface of IoT architectures often suffers from compromised security paradigms, rendering smart environments highly vulnerable to malicious exploitations. While traditional Intrusion Detection Systems (IDS) mitigate network threats, conventional datasets lack the granular, protocol-specific traffic anomalies characteristic of IoT environments. This research addresses this gap by developing an automated machine learning framework designed to differentiate reconnaissance and anomalous activities from baseline behaviors within smart home IoT infrastructures. Utilizing the Hacking and Countermeasure Research Lab (HCRL) dataset, we evaluate and contrast the efficacy of Naïve Bayes (NB) and Support Vector Machine (SVM) algorithms across varying data-split ratios. Experimental results indicate that while Naïve Bayes offers competitive computational recall in localized environments, the SVM classifier demonstrates superior robustness, achieving an accuracy threshold approaching 99.99% in isolating low-frequency reconnaissance attacks.
A machine learning-based framework to tackle issues in traditional systems in traditional systems is introduced by combining large language models (LLMs) and is effective in identifying possible threats as well as filling the semantic gap.
Mamoon M. Saeed, Rashid A. Saeed, Salah Hagahmoodi et al.· Baghdad Science Journal· 0 citations
The rapid proliferation of Internet of Things (IoT) devices and their integration into increasingly interconnected applications have substantially expanded the attack surface of modern networked systems. The heterogeneous nature and high volume of IoT traffic make timely and reliable identification of malicious activities increasingly important for maintaining the security and resilience of IoT-enabled environments. This study investigates the effectiveness of maching learning approaches for supervised malicious traffic classification in IoT networks using the ACI-IoT-2023 dataset. A comparative experimental study is conducted across binary and eleven-class classification tasks to examine the capability of different learning approaches to distinguish benign and malicious traffic and identify diverse attack categories. The results demonstrate strong classification performance across the evaluated approaches, with XGBoost achieving the highest ROC-AUC in binary classification and the Decision Tree delivering the best overall performance in eleven-class classification. Further analysis of feature importance identifies several flow-level features that contribute substantially to classification performance. Overall, the findings demonstrate the effectiveness of machine learning-based approaches for accurate and efficient malicious traffic classification in IoT networks.
Connor Gladish, Molly Corgan, J. Moss et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.