Aug 2026· Mathematical Modeling and Algorithm Application· 0 citations· 28 references
TL;DR
This study confirms that FL is a feasible method for distributed air quality prediction, and the key point is to establish a prediction model, not to change the algorithm.
Abstract
Ground ozone and other air pollutants pose a great threat to the health. Because the data of different monitoring stations are scattered, it is very complicated to accurately predict the ozone level. Federated Learning (FL) allows everyone to train models together without exchanging raw data, thus protecting privacy. This paper uses FL to predict air quality, and the key point is to establish a prediction model, not to change the algorithm. This paper used UCI air quality data set (with 8,762 records, each with four attributes) to simulate the situation of five users, and the data were separated by independent identically distributed (IID) and non-independent identically distributed (Non-IID). This paper trained 50 communication rounds with a multi-layer perceptron (MLP) and Federal Average (FedAvg) method. The model of centralized training and local training is also used to compare. The experimental results show that the average absolute error (MAE) of FedAvg is 48.9 under IID data and 58.3 under Non-IID data, which is much better than the local training models (MAE 78.5 and 85.2) and close to the effect of centralized training (MAE 42.3 and 51.6). This method of federated learning adapts well to different data. This study confirms that FL is a feasible method for distributed air quality prediction.
Empirical data on the impact of data heterogeneity on federated learning is provided, and it is proved that federated learning can be a viable alternative in privacy-sensitive environmental prediction problems.
Zhe-Yu Qiu· Mathematical Modeling and Al...· 0 citations
The core objective of this research is to construct a machine learning-based air quality prediction model. This model aims to forecast the Air Quality Index (AQI) for the next 72 hours and classify its corresponding levels (e.g., Good, Moderate, Polluted), providing a robust scientific basis for environmental protection departments and related decision-making. For feature selection, we analyzed multiple key factors affecting air quality. While meteorological data, spatiotemporal features, and external pollution sources are important, this study focuses on the historical concentrations of six critical pollutants (PM2.5, PM10, SO2, NO2, CO, and O3) as model inputs to establish a baseline model, acknowledging the need for incorporating broader influencing factors in future work. In the model construction phase, we performed extensive preprocessing on the collected historical air quality data, including standardization and normalization, to extract effective information. We then employed and compared several advanced machine learning algorithms, selecting the optimal combination to build the final prediction model. The experiments were conducted using the Python language. By continuously optimizing model parameters and feature combinations, we achieved predictions for both the numerical AQI values and their corresponding quality levels for the subsequent 72 hours. Experimental results demonstrate that the constructed model possesses high prediction accuracy and stability for the predominant “Excellent” and “Good” categories. However, the lack of severe pollution events in the dataset limits the evaluation of its predictive capability for pollution peak events.
Shiting Wu, Xiaohua Qian, Xiaodong Zhou et al.· 2026 3rd World Conference on...· 0 citations
There are significant gaps that remain in terms of model interpretability and the ability to generalize across climate variations, so this article provides a relatively comprehensive overview of the application of federated learning in air quality forecasting and monitoring.
Yuhao Wu· Mathematical Modeling and Al...· 0 citations
Air quality forecasting models anticipate and control pollution concentrations quickly. In this research, two diversified datasets are analyzed and modeled using the proposed novel Hybrid Forecasting Model (HFM). PM10, PM2.5, NO, and NO2 pollutant concentrations for the National Capital Region, Delhi, were collected from the Central Pollution Control Board as Dataset 1. A Raspberry Pi hardware integrated with an MCP3008 Analog-to-Digital Converter and pollutant sensors like MQ135 and GP2Y1010AU0F was assembled to collect PM2.5, CO, and NH3 pollutant concentrations at the study site as Dataset 2. The results obtained for the proposed HFM are compared with three existing algorithms in terms of R-squared and Mean Squared Log Error. The novelty of this work is to test the performance of the algorithm by executing it on a computer with an Intel 8265U processor and also on a Raspberry Pi integrated with an MCP3008 Analog-to-Digital Converter. The experimental results indicate that the Raspberry Pi requires approximately 3.37% of the power consumed by the Intel 8265U processor. Conversely, the Intel 8265U processor achieves approximately 86.67% lower latency, thereby delivering significantly faster computational performance compared to the Raspberry Pi.
D. Subramanian, P. Marimuthu, Nithyalakshmi Ramadoss· Journal of Trends in Compute...· 1 citation
Air pollution has become a major environmental and public-health concern worldwide, and understanding its behaviour is essential for effective monitoring and management. This study investigates air-quality patterns across four regions in the Sultanate of Oman—Al Khuwair, Salalah, Al Khoud, and Bediya—using a combination of statistical modelling and machine-learning techniques. Hourly data for 2023, including pollutant concentrations and key meteorological variables, were obtained from the Environment Authority of Oman, cleaned, and pre-processed to construct region-specific datasets. Air Quality Index (AQI) values were calculated for each pollutant and classified into three categories (Good, Moderate, and Unhealthy). Kernel Support Vector Machine (KSVM) and Gaussian Process Regression and models were trained using a 70/30 temporal split to classify AQI levels. Results showed that KSVM achieved the highest accuracy in Salalah (96.97%), Al Khoud (94.33%), and Bediya (93.37%), while Gaussian Process Regression performed best in Al Khuwair (70.32%). In conclusion, this research demonstrates that advanced kernel-based classifiers can effectively model non-linear environmental data, providing a scalable solution for regional environmental management.
Shamssa Abdullah Al-Rahbi, M. Alodat· SISTEMASI· 0 citations
An optimized machine learning-based Air Quality Forecasting System that integrates Extreme Learning Machines (ELM) and Genetic Algorithms (GA) to predict short-term variations in air quality and demonstrates a robust, scalable, and practical solution for short-term air quality prediction.
Shivatejaswini B, Kiran B. M., G. Prasad· International Scientific Jou...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.