Assessment of machine learning methods for urban air pollution forecasting
Abstract
Fine particulate matter (PM2.5) forecasting supports public-health advisories and operational early-warning systems. We present a European, multi-country benchmark for direct, multi-horizon PM2.5 forecasting (1/3/6/12/24 h) that compares statistical, tabular machine learning, and sequence deep learning models under a single, reproducible experimental design. We construct a harmonized hourly dataset (2018–2024) by joining the European Environment Agency (EEA) station measurements with meteorology and station metadata and evaluate two complementary protocols: Protocol A—maximum tabular coverage—and Protocol B—a common sequence-eligible subset enabling cross-paradigm fairness. Across horizons, boosted trees (LightGBM/XGBoost) are consistently strong under Protocol A, while under Protocol B residual long short-term memory (LSTM) variants (with attention at h = 1) are competitive at short horizons and boosted trees dominate at medium–long horizons. Stratified analyses reveal substantial heterogeneity by country and station area, motivating horizon-specific models and stratified monitoring in deployment. We further quantify the coverage–comparability trade-off, report skill vs persistence, and paired significance tests, and provide feature-importance summaries to aid interpretation. All code, configurations, and masks are released for full reproducibility, establishing a transparent baseline for future methodological advances.