Aug 2026· Dasinya Journal for Engineering and Informatics· 0 citations· 25 references
TL;DR
A Temporal-Aware Multi-Task Learning (TMTL-AQI) framework to assess urban air quality via structured data that outperforms single-task and baseline multi-task models with an accuracy of 0.7605 and an F1-score of 0.7422.
Abstract
Accurate air pollution forecasting is vital for protecting the environment and public health. However, predicting air pollution continues to be a challenge due to complicated relationships between meteorological variables, pollutant concentrations, and time-dependent characteristics. This paper proposes a Temporal-Aware Multi-Task Learning (TMTL-AQI) framework to assess urban air quality via structured data. The model simultaneously performs three tasks: Air Quality Index category classification, PM2.5 and PM10 particulate matter regression, and auxiliary AQI value prediction. The framework consists of data preprocessing, cyclical temporal encoding, a shared backbone based on hard parameter sharing, and three task-specific output heads optimized jointly. It uses cyclical transformations to encode temporal traits that capture periodic patterns of the environment and produces a common neural representation through hard parameter sharing. The TRAQID dataset was used for empirical testing. The model outperforms single-task and baseline multi-task models with an accuracy of 0.7605 and an F1-score of 0.7422 for AQI classification. Although based solely on structured input (no image-based features), the model performed competitively with a state-of-the-art approach due to the reduced prediction error in regression tasks (MAE values of 13.34 μg/m³ for PM2.5 and 22.06 μg/m³ for PM10).
An optimized machine learning-based Air Quality Forecasting System that integrates Extreme Learning Machines (ELM) and Genetic Algorithms (GA) to predict short-term variations in air quality and demonstrates a robust, scalable, and practical solution for short-term air quality prediction.
Shivatejaswini B, Kiran B. M., G. Prasad· International Scientific Jou...· 0 citations
Air pollution has become one of the most critical environmental and public health challenges in India. Major metropolitan cities such as Delhi, Mumbai, and Hyderabad frequently record hazardous Air Quality Index (AQI) levels due to vehicular emissions, industrial activities, meteorological variability, and urbanization. Accurate short-term and multi-horizon AQI forecasting is essential for early warning systems and policy intervention. This study proposes a comprehensive spatio-temporal deep learning framework integrating Long Short-Term Memory (LSTM) networks and Graph Neural Networks (GNN) for multi-city AQI prediction using Central Pollution Control Board (CPCB) data from 2018–2024. The model captures both temporal pollutant dynamics and spatial inter-city correlations. Comparative evaluation against ARIMA and Random Forest models demonstrates that the proposed hybrid model reduces RMSE by 18–25% for 1-hour forecasting and 15–20% for 24-hour forecasting. Statistical significance testing confirms robustness (p < 0.05). The results indicate that spatial dependency modeling significantly enhances predictive performance in urban air quality systems.
Dr. Mohammed Sharfuddin, Asma Fatima· International Journal of Eng...· 0 citations
Air quality forecasting has become an important tool for the management of public health and urban planning with the rapid increase of the urban population and the vehicular traffic. Although deep learning models have been able to predict the Air Quality Index (AQI) with a high accuracy, their black-box nature limits their use by regulators and city planners who need transparent and trustworthy support for their decision-making. This paper aims to design a explainable deep learning model taking into account meteorological parameters, traffic-flow factors and historical pollutant information to predict the AQI and detect abnormal pollution events. The combination of CNN-BiLSTM with an attention mechanism is used to model the spatial correlation between pollutants and long-term temporal correlations. To understand the model, SHapley Additive exPlanations (SHAP) and Integrated Gradients, which quantify the contribution of different meteorological and traffic features to each prediction, are used. An unsupervised residual-thresholding module also identifies abnormal AQI episodes that do not match the normal pattern, e.g. due to an increase in traffic congestion or an atmospheric temperature inversion. The experimental results using a multi-source dataset that is a combination of air-quality monitoring records, meteorological observations, and traffic-sensor data demonstrate that the proposed model achieves a coefficient of determination (R²) of 0.943 and a root-mean-square error (RMSE) of 9.82, surpassing the baseline LSTM, GRU, and CNN-LSTM models. The lagged PM2.5 concentration, traffic volume and relative humidity are the top three factors, which agree with the knowledge of the atmospheric chemistry, and thus confirms the validity of the model's reasoning process. The proposed framework shows that the prediction accuracy and interpretability can be achieved concurrently, and can provide a practical and transparent decision support tool for environmental agencies and smart-city traffic management systems.
Meena Kumari, Jitendra Nath Shrivastava, Arijit Dey· International journal of com...· 0 citations
AirFlow is proposed, a pollutant-aware dual-stream framework that operates on station multivariate observations without additional graph propagation or predefined signal decomposition, achieving high forecasting accuracy with low computational overhead.
Fine particulate matter (PM2.5) forecasting supports public-health advisories and operational early-warning systems. We present a European, multi-country benchmark for direct, multi-horizon PM2.5 forecasting (1/3/6/12/24 h) that compares statistical, tabular machine learning, and sequence deep learning models under a single, reproducible experimental design. We construct a harmonized hourly dataset (2018–2024) by joining the European Environment Agency (EEA) station measurements with meteorology and station metadata and evaluate two complementary protocols: Protocol A—maximum tabular coverage—and Protocol B—a common sequence-eligible subset enabling cross-paradigm fairness. Across horizons, boosted trees (LightGBM/XGBoost) are consistently strong under Protocol A, while under Protocol B residual long short-term memory (LSTM) variants (with attention at h = 1) are competitive at short horizons and boosted trees dominate at medium–long horizons. Stratified analyses reveal substantial heterogeneity by country and station area, motivating horizon-specific models and stratified monitoring in deployment. We further quantify the coverage–comparability trade-off, report skill vs persistence, and paired significance tests, and provide feature-importance summaries to aid interpretation. All code, configurations, and masks are released for full reproducibility, establishing a transparent baseline for future methodological advances.
Bianca-Iuliana Chisilev, Elena Pelican· Stochastic environmental res...· 0 citations
The core objective of this research is to construct a machine learning-based air quality prediction model. This model aims to forecast the Air Quality Index (AQI) for the next 72 hours and classify its corresponding levels (e.g., Good, Moderate, Polluted), providing a robust scientific basis for environmental protection departments and related decision-making. For feature selection, we analyzed multiple key factors affecting air quality. While meteorological data, spatiotemporal features, and external pollution sources are important, this study focuses on the historical concentrations of six critical pollutants (PM2.5, PM10, SO2, NO2, CO, and O3) as model inputs to establish a baseline model, acknowledging the need for incorporating broader influencing factors in future work. In the model construction phase, we performed extensive preprocessing on the collected historical air quality data, including standardization and normalization, to extract effective information. We then employed and compared several advanced machine learning algorithms, selecting the optimal combination to build the final prediction model. The experiments were conducted using the Python language. By continuously optimizing model parameters and feature combinations, we achieved predictions for both the numerical AQI values and their corresponding quality levels for the subsequent 72 hours. Experimental results demonstrate that the constructed model possesses high prediction accuracy and stability for the predominant “Excellent” and “Good” categories. However, the lack of severe pollution events in the dataset limits the evaluation of its predictive capability for pollution peak events.
Shiting Wu, Xiaohua Qian, Xiaodong Zhou et al.· 2026 3rd World Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.