Skip to content

Empowering Non-IID Federated Learning With Data Augmentation and Data-Free Knowledge Distillation

2026 · IEEE Transactions on Information Forensics and Security · Vol 21, pp. 6420-6435 · 0 citations · 43 references
Computer Science

Abstract

Federated learning (FL) is an emerging distributed machine learning framework that enables collaborative learning among multiple parties while preserving data privacy. However, the complexity of environments and node heterogeneity in the real world result in uneven data distribution across nodes, leading to Non-IID (Non-Independent and Identically Distributed) characteristics in data distribution. Such data distribution significantly reduces the convergence and performance of the model, becoming one of the fundamental challenges in federated learning mechanisms. To address the above issue, this paper proposes a novel FL framework-FedGKD. For the Non-IID client data distribution problem, we employ client-side local data augmentation, where GAN models are deployed on each client to generate synthetic samples so that local data distribution imbalance can be effectively alleviated. To further overcome the limitations of client-side local data augmentation under Non-IID, FedGKD introduces server-side privacy-preserving data-free knowledge distillation, which can transfer the knowledge of selected clients to the server while ensuring privacy protection, further mitigating the impact of Non-IID on federated learning and solving the problem of model performance degradation caused by direct aggregation. Extensive experiments demonstrate that FedGKD significantly outperforms the baseline algorithms in terms of accuracy, while exhibiting excellent performance in other metrics.

View source

Similar papers

Conference Jul 2026

Privacy-Preserving Federated Learning Framework for Robust Model Training Under Non IID Data Distributions

Federated learning is a decentralised machine-learning approach in which several clients jointly build a shared model without moving their raw data to one location. Rising concerns around privacy, tightening regulation, and restrictions on how data may be owned or shared have made this approach increasingly attractive in practice. Although federated learning lowers privacy exposure relative to centralised training, deploying it in practice is complicated by clients whose data are unevenly distributed and non-identically distributed, by clients that participate inconsistently, and by training that can converge unpredictably. To obtain global models that train reliably and consistently even when client data are heterogeneous, this work puts forward a federated learning system built around privacy preservation. The design follows a client–server pattern in which a coordinating server aggregates updates from local models using weights that account for imbalance among participants. The behaviour of the resulting system is examined methodically across several data-distribution regimes — IID, mildly non-IID, and severely non-IID. The experiments show that the framework converges reliably and delivers predictive accuracy that holds up well, especially in the more difficult non-IID cases. Compared with conventional federated learning baselines, the approach shows greater robustness and steadier performance across successive training rounds. Because it is simple to implement, repeatable, and built with real deployment in mind, the architecture suits privacy-sensitive, decentralised use cases such as distributed intelligent systems, industrial monitoring, and healthcare analytics.

Shyam Patel, S. Khan · 0 citations
Open access Jul 2026

Personalized Data-Free Knowledge Distillation for Federated Learning under Heterogeneous Models and Data

PDKD solves the problem of model drift caused by the inconsistent distribution of distillation datasets and the local data by generating personalized distillation datasets for each client while protecting client data privacy.

Jing-feng Tu, Lei Yang, Chao Ma et al. · 0 citations
Conference Open access Jul 2026

Federated Learning with Differential Privacy: A Comprehensive Framework for Privacy-Preserving Distributed Machine Learning

This study implemented a comprehensive experimental framework for analysing FL performance using standard FL aggregation protocols FedAvg, FedProx, and SCAFFOLD in conjunction with Differential Privacy mechanisms; specifically, the Gaussian noise mechanism with Rényi Differential Privacy (RDP) accountants.

Himanshi Singh, Kahksha Ahmed, Priyanshu Prajapati et al. · 0 citations
Preprint Aug 2026

Out-of-Distribution Federated Distillation with Domain-Aware Proxy

A domain-aware proxy selection framework to better adopt proxy data for OOD problems is proposed and the experimental results show that the proposed models effectively address the challenges of distribution shifts under OOD with and without proxy data.

Jiahao Xiao, Jiangming Liu · 0 citations
Aug 2026

EFFEKT: Efficient Federated Knowledge Transfer to Foundation Models

This work introduces a novel multi-domain federated learning framework in which lightweight client-side proxy models collaborate with a server-side Foundation Model (FM) to learn new concepts without sharing private data.

Matteo Caligiuri, Francesco Barbato, Pietro Zanuttigh et al. · 0 citations
Jul 2026

Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning

FedKT-CSD (Federated Knowledge Transfer via Collaborative Synthetic Data), a framework inspired by neural image compression that closes the gap in jointly achieving low communication, robustness to heterogeneity, and rigorous privacy by leveraging publicly pretrained autoencoders as a shared latent space.

Maximilian Andreas Hoefler, Karsten Mueller, Wojciech Samek · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.