Skip to content
Open access

On the Impact of Data Heterogeneity in Federated Learning: A Case Study on Air Quality Prediction

Aug 2026 · Mathematical Modeling and Algorithm Application · 0 citations · 12 references

TL;DR

Empirical data on the impact of data heterogeneity on federated learning is provided, and it is proved that federated learning can be a viable alternative in privacy-sensitive environmental prediction problems.

Abstract

Federated learning has become a popular model to apply in privacy-preserving modeling in distributed settings, particularly when the models are applied to data that is distributed among various locations and is often sensitive, such as in air quality prediction. This paper looks into how effective federated learning is for predicting ozone (O₃) concentrations, under both independent and non-independent data distributions. In particular, two representative algorithms Federated Averaging (FedAvg)and Federated Averaging (FedAvg) are experimented on a real-world air quality dataset. An experimental framework was established that was relatively comprehensive and centralized training and local-only models were used as baselines. Mean absolute error (MAE) and root mean squared error (RMSE) are used to measure model performance. The findings suggest that federated learning performs much better than the isolated local models and the performance is similar to that of the centralized training. Having said that, data heterogeneity does present certain issues-it slows down convergence and decreases accuracy in prediction. In such non-IID conditions, FedProx is more stable and less erroneous than FedAvg implying that it is more resistant to client drift.Most importantly, this paper provides empirical data on the impact of data heterogeneity on federated learning, and proves that federated learning can be a viable alternative in privacy-sensitive environmental prediction problems.

Read PDF

Similar papers

Review Open access Aug 2026

An Investigation of Federated Learning for Air Quality Forecasting and Monitoring

There are significant gaps that remain in terms of model interpretability and the ability to generalize across climate variations, so this article provides a relatively comprehensive overview of the application of federated learning in air quality forecasting and monitoring.

Yuhao Wu · 0 citations
Open access Aug 2026

The Impact of Data Preprocessing on Federated Learning for Air Quality Prediction: A Comparative Study of Imputation Methods

Data preprocessing plays a foundational role in machine learning, but it receives limited systematic attention in federated learning (FL) environments. This study empirically compares four missing value strategies using the UCI dataset: direct deletion, mean imputation, K-Nearest Neighbors (KNN) imputation, and Random Forest (RF) imputation. The setup uses a federated framework with five clients, each running a multi layer perceptron model. The findings show direct deletion achieves the strongest performance, with a mean absolute error of 0.173 and an  of 0.966, clearly exceeding imputation methods (the errors around 0.25). The NMHC(GT) feature has an 88.4% missing rate, making imputation unreliable and introducing noise. Although direct deletion reduces the sample from 7,674 to 827 observations, it safeguards data integrity. This research shows that preserving data quality should take priority over quantity when the missing rates are exceptionally high, providing the practical guidance for robust preprocessing design in federated learning systems.

Xiao Liu · 0 citations
Conference Jul 2026

Privacy-Preserving Crop Disease Prediction using Federated Meta-Learning and Climate Intelligence

The growing diversity of climate conditions and the scale effects of agricultural data pose serious problems for the accuracy and scalability of crop disease prediction. The drawbacks of traditional centralized machine learning methods include limitations caused by data privacy, low generalization, and high communication overheads. To overcome these challenges, in this study, we propose a new federated meta-learning with climate-driven personalization (FML-CDP) system that incorporates federated learning, meta-learning, and climate-sensitive modelling in a single architecture. The suggested system allows decentralized and distributed training to several farms without losing data privacy and addresses non-identically distributed (non-IID) data. The framework is more accurate in terms of predictions and more context-aware by using multimodal inputs, such as leaf images, IoT sensor data, and climate variables. The meta-learning aspect allows for quick local farm adjustments and individual predictions. The proposed model outperformed the baseline approach, achieving an accuracy of 97.3% and an F1-score of 0.96, surpassing the performance of centralised convolutional neural networks (CNN), FedAvg, and meta-learning models. Another benefit of climate intelligence is that it enhances resilience and flexibility of the system. The proposed framework offers privacy-saving, scalable, and efficient next-generation smart agricultural systems.

Mary Navyatha Govindu, Chiranjeevi Manike · 0 citations
Open access Jul 2026

Privacy-preserving load forecasting in smart grids using federated learning: a comparative analysis of aggregation strategies

Experimental results on real-world energy consumption datasets demonstrate that the proposed FL framework achieves competitive forecasting accuracy while preserving client data privacy, and a rigorous comparative analysis reveals that FedProx and FedTrimmedAvg consistently outperform FedAvg under non-IID conditions.

A. Tibermacine, Ilyes Naidji, Imad Eddine Tibermacine et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.