Skip to content
Open access

Privacy-Preserving Feature Engineering for Federated Learning Analytics

2021 · International Journal of Data Engineering and Intelligent Computing · 0 citations

TL;DR

This paper proposes a comprehensive framework for privacy-preserving feature engineering (PPFE) within federated learning analytics and explores techniques such as homomorphic encryption, differential privacy, and secure multi-party computation to enable robust, privacy-safe feature selection, transformation, and extraction across clients.

Abstract

Federated learning (FL) offers a decentralized approach to machine learning that preserves data privacy by training models locally across distributed devices. However, the feature engineering process—an essential step in improving model performance—often remains centralized or privacy-invasive, risking sensitive data exposure. This paper proposes a comprehensive framework for privacy-preserving feature engineering (PPFE) within federated learning analytics. We explore techniques such as homomorphic encryption, differential privacy, and secure multi-party computation to enable robust, privacy-safe feature selection, transformation, and extraction across clients. Our framework includes both vertical and horizontal FL settings and evaluates the trade-offs between privacy, utility, and communication overhead. Experimental results on real-world datasets demonstrate that our PPFE methods can significantly improve model performance without compromising data privacy. This work contributes towards building a more secure and efficient FL pipeline that ensures end-to-end data confidentiality.

Read PDF

Similar papers

Open access 2018

A Framework for Privacy-Preserving Machine Learning in Sovereign Cloud Environments

A novel framework for integrating privacy-preserving ML techniques within Sovereign Cloud infrastructures is presented, combining cutting-edge cryptographic approaches with the data sovereignty features of Sovereign Clouds, ensuring data privacy, legal compliance, and efficient machine learning at scale.

Ahmed Hassan · 0 citations
Open access 2026

Convex Optimization Methods for Privacy-Preserving Federated Learning

This work provides a unified approach for aiding the design of state-of-the-art privacy-preserving distributed learning systems that are also utility-optimal and is an important step towards using such approaches in high-stakes domains like healthcare or finance.

A. M., Nitish Kumar · 0 citations
2026

Differential Privacy Enabled Cascaded Filter for Efficient and Privacy-Preserving Federated Learning

Federated learning (FL) enables collaborative model training across multiple clients in a privacy-preserving manner. However, the employment of homomorphic encryption algorithms might lead to high computational cost while the application of differential privacy (DP) methods would sacrifice model performance. To establish efficient and secure FL system as well as maintaining competitive performance, we introduce a DP-enabled cascaded filter with novel model-aggregation mechanism. Taking the model parameters of one client for example, dimensions with large absolute values and significant variations are selected by the cascaded filter and regarded as important dimensions. After that, random noise are added to these important dimensions for data security. Finally, the model parameters will be sent to the central server for aggregation. Unlike traditional DP-based approaches, our method considers each dimension’s informational importance, i.e., both absolute value and its variation. We theoretically prove the convergence of our method and verify the effectiveness on different scenarios including four datasets. The experimental results suggest that our method outperforms the other ones from literature under both IID and non-IID conditions while achieving high level of efficiency and privacy protection performance.

Zhiqiang Chen, Yuchen Jiang, Ray Y. Zhong et al. · 0 citations
Conference Jul 2026

NETFL: Network Efficient Fully Secure Federated Learning at Scale

Federated Learning (FL) enables collaborative model training without centralizing client data, making it well-suited for privacy-sensitive domains. Existing approaches use techniques such as homomorphic encryption, differential privacy, and secure multi-party computation to mitigate attacks including model inversion, membership inference, and gradient leakage. However, these methods often incur high computational and memory overheads and frequently overlook confidentiality of the global model itself, which may be proprietary and sensitive. These limitations reduce the practicality of secure FL in large-scale and compliance-sensitive environments.We present NETFL, a Fully Secure and scalable FL framework that decentralizes training across client pairs using lightweight MPC, while restricting servers to secure aggregation, client pairing, and routing. NETFL eliminates server-side bottlenecks, avoids full data offloading, and preserves confidentiality of data, model parameters, and updates throughout training. Our evaluation shows that NETFL protects against unauthorized observation, reconstruction, gradient leakage, membership inference, and inversion attacks, while achieving up to 13× faster training time and 50% lower server memory usage compared to prior work.

Sahar Ghoflsaz Ghinani, Elaheh Sadredini · 0 citations
Open access 2026

A Study on Federated Learning Techniques for Privacy Preservation

An in-depth analysis of federated learning methods and paying special attention to the issue of privacy is provided, which examines new developments, concerns and tradeoffs connected with privacy, effectiveness of communication, model noise, and scalability of systems.

Aarav Mehta · 0 citations
#federated learning Open access Sep 2026

Trustworthy and privacy-preserving decentralized federated learning with multi-layer defense for secure collaborative AI

Decentralized Federated Learning (DFL) enables collaborative artificial intelligence model training without centralizing sensitive data, making it suitable for privacy-critical and distributed intelligent systems such as healthcare, Industrial IoT, and smart digital infrastructure. Despite its advantages, DFL remains vulnerable to privacy leakage through shared model updates and to model poisoning and backdoor attacks that compromise system reliability, robustness, and trustworthiness. Existing defense mechanisms primarily address either privacy preservation or poisoning robustness independently and often exhibit limited effectiveness under adaptive or high-ratio adversarial settings. This work proposes a trustworthy and privacy-preserving decentralized federated learning framework that jointly addresses these challenges through two integrated components: (i) a hybrid privacy mechanism based on public dataset pretraining followed by differentially private fine-tuning, and (ii) a multi-layer model defense architecture designed to mitigate poisoning and backdoor attacks across decentralized peer-to-peer environments. The framework integrates local data sanitization, peer-side model verification, robust trimmed-mean aggregation, and runtime inference protection to provide defense-in-depth across both training-time and inference-time attack surfaces. An adversary model and operational assumptions are formally defined, and the framework is evaluated under strong adversarial conditions, including a 20% poisoning ratio. Experimental results demonstrate consistent robustness improvements over a vanilla DFL baseline. While the baseline model achieves a clean accuracy of 83.10%, the proposed framework improves clean performance to 86.12%. Under adversarial conditions, accuracy improves from 37.71% to 53.88% for Fast Gradient Sign Method (FGSM) attacks, from 21.75% to 46.40% for Projected Gradient Descent (PGD) attacks, and from 40.62% to 67.35% for Carlini–Wagner (CW) attacks. For backdoor-based poisoning attacks such as BadNets and Blended attacks, the defense pipeline restores model accuracy to above 86% while maintaining stable benign performance. These findings demonstrate that the proposed framework provides an effective balance between privacy preservation, adversarial robustness, and trustworthy decentralized collaborative learning for secure AI-driven systems.

Durga Sivan, Uma Maheshwari Shanmugam, Sachnev Vasily et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.