Skip to content
Conference

A Machine Learning-Based Model for 72-Hour Air Quality Prediction and Classification

Jul 2026 · 2026 3rd World Conference on Computer and Information Security (WCCIS) · pp. 40-45 · 0 citations · 38 references

Abstract

The core objective of this research is to construct a machine learning-based air quality prediction model. This model aims to forecast the Air Quality Index (AQI) for the next 72 hours and classify its corresponding levels (e.g., Good, Moderate, Polluted), providing a robust scientific basis for environmental protection departments and related decision-making. For feature selection, we analyzed multiple key factors affecting air quality. While meteorological data, spatiotemporal features, and external pollution sources are important, this study focuses on the historical concentrations of six critical pollutants (PM2.5, PM10, SO2, NO2, CO, and O3) as model inputs to establish a baseline model, acknowledging the need for incorporating broader influencing factors in future work. In the model construction phase, we performed extensive preprocessing on the collected historical air quality data, including standardization and normalization, to extract effective information. We then employed and compared several advanced machine learning algorithms, selecting the optimal combination to build the final prediction model. The experiments were conducted using the Python language. By continuously optimizing model parameters and feature combinations, we achieved predictions for both the numerical AQI values and their corresponding quality levels for the subsequent 72 hours. Experimental results demonstrate that the constructed model possesses high prediction accuracy and stability for the predominant “Excellent” and “Good” categories. However, the lack of severe pollution events in the dataset limits the evaluation of its predictive capability for pollution peak events.

View source

Similar papers

Open access Aug 2026

Optimized and Explainable Air Quality Index Classification System

Air pollution has become a major environmental and public health concern due to rapid urbanization, industrial growth, and increasing vehicular emissions. High concentrations of pollutants such as PM2.5, PM10, NO₂, SO₂, CO, and O₃ can significantly impact human health and environmental sustainability. Accurate monitoring and prediction of air quality are therefore essential for effective environmental management and public safety. This paper presents AirAware, a machine learning–based system designed to predict and monitor Air Quality Index (AQI) levels using historical air pollution data and real-time environmental information. The system utilizes the XGBoost algorithm to analyze pollutant parameters and generate accurate AQI predictions and classifications. Data preprocessing techniques such as cleaning, normalization, and SMOTE-based class balancing are applied to improve model performance and ensure reliable predictions across different AQI categories. In addition, the system integrates real-time air pollution data through the OpenWeather API, enabling continuous monitoring of current environmental conditions. The predicted AQI values and pollution trends are displayed through a web-based dashboard, allowing users to visualize air quality patterns and compare real-time data with machine learning predictions. By combining machine learning techniques with real-time data integration, the proposed system provides an effective solution for air quality prediction, monitoring, and environmental awareness.

Neethu Roy, Jeeson Justin · 0 citations
Open access Jul 2026

Prediction of Air Pollution in the Sultanate of Oman using Machine Learning Approaches

Air pollution has become a major environmental and public-health concern worldwide, and understanding its behaviour is essential for effective monitoring and management. This study investigates air-quality patterns across four regions in the Sultanate of Oman—Al Khuwair, Salalah, Al Khoud, and Bediya—using a combination of statistical modelling and machine-learning techniques. Hourly data for 2023, including pollutant concentrations and key meteorological variables, were obtained from the Environment Authority of Oman, cleaned, and pre-processed to construct region-specific datasets. Air Quality Index (AQI) values were calculated for each pollutant and classified into three categories (Good, Moderate, and Unhealthy). Kernel Support Vector Machine (KSVM) and Gaussian Process Regression and models were trained using a 70/30 temporal split to classify AQI levels. Results showed that KSVM achieved the highest accuracy in Salalah (96.97%), Al Khoud (94.33%), and Bediya (93.37%), while Gaussian Process Regression performed best in Al Khuwair (70.32%). In conclusion, this research demonstrates that advanced kernel-based classifiers can effectively model non-linear environmental data, providing a scalable solution for regional environmental management.

Shamssa Abdullah Al-Rahbi, M. Alodat · 0 citations
Open access Aug 2026

Artificial Intelligence–Driven Solutions in Environmental Engineering: Advancing Sustainable Urban Air Quality Through PM2.5 Prediction Using Machine Learning

Air pollution prediction is an integral component of environmental engineering, and effective management of air quality demands timely, accurate and explainable air quality forecasting to manage urban pollution. The concept of this research was to use artificial intelligence to predict the PM2.5 level using the required predictors PM10 and NO₂ as well as other pollutant, meteorological, temporal, wind-direction and site-specific variables from the Beijing Multi-Site Air Quality dataset. A quantitative supervised regression approach was followed, covering data preprocessing techniques such as data cleaning, handling of missing values, temporal feature engineering, encoding of categorical variables, selection of models, hyperparameter tuning, cross-validation, residual analysis, error analysis for various categories of pollution, feature importance analysis, and explainability using SHAP. A common data mining pipeline was used to evaluate the performance of five models (Linear Regression, Decision Tree, Random Forest, Gradient Boosting, XGBoost). Random Forest Regressor achieved maximum predictive performance in terms of high accuracy, low validation error, low residual bias and acceptable performance in all the pollution categories. The importance of PM10 was identified by the explainability analysis, while the other environmental factors also provided valuable information. The results indicate that explainable ensemble learning can provide a solid basis for AI-based PM2.5 prediction and assist in air-quality early warning, environmental monitoring, pollution-control and sustainable urban management.

C. V. S. S. P. Kumar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.