2026· IEEE Journal on Selected Areas in Communications· Vol 44, pp. 5712-5729· 0 citations· 58 references
Abstract
Fine-tuning remains essential for adapting large models to diverse downstream tasks, yet doing so in a privacy-preserving and resource-efficient manner is challenging, particularly in federated learning (FL) on edge devices. Parameter tensor freezing is a promising solution. However, current methods face key limitations. Static tensor freezing struggles to adapt to the non-IID (non-independent and identically distributed) data distributions across FL clients, while localized dynamic freezing may lead to slow convergence or divergence across clients, harming overall accuracy. We propose FedFreeze, a communication-aware and dynamic tensor-freezing framework for federated fine-tuning over resource-constrained edge networks. FedFreeze delivers two key benefits: 1) it explicitly incorporates computation costs and bandwidth-dependent communication costs into freezing-mask optimization, actively selecting which tensors are updated and transmitted to reduce computation and communication overheads; and 2) it improves convergence stability by coordinating freezing decisions based on sampled client statistics. We further analyze the memory usage patterns of FedFreeze and introduce the first kind of memory management strategy to minimize memory consumption of tensor-freezing based methods in FL. Experimental results based on real-world traces from NVIDIA Jetson hardware demonstrate that FedFreeze accelerates convergence by up to $5.46\times $ and reduces peak memory usage by up to 53.9% without compromising model accuracy. Furthermore, evaluations under heterogeneous computation and communication environments confirm its robustness.
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental"aggregation dilemma"between the accurate Sum-of-Products (SoP) and the communic...
Han Zou, Chao Zhang, Yu-Zhi Yang et al.· 0 citations
SplitLite is proposed, a communication-efficient split federated LoRA fine-tuning method that exploits the low effective rank structure of consecutive-epoch activation and gradient residuals, thereby significantly reducing both activation uplink and gradient downlink traffic.
Energy-Aware Adaptive Quantization and Freezing (EA-AQF), a unified framework that co-optimizes communication and computation, is presented, a unified framework that co-optimizes communication and computation and maintains robust convergence in highly heterogeneous tasks.
Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure fo...
Meng-Jun Yi, Huai-An Gu, Yi-Hao Ai et al.· 0 citations
FCCA, which estimates the signed input--error cross-covariance, whitens it with diagonal Fisher moments, truncates it in the resulting local metric, maps the selected directions back, and applies thin QR to obtain stable core coordinates, shows that a carefully selected fixed span can recover most of the benefit of mov...
Wen-Song Ye, Zhan-Ming Shen, Zhiqing Xiao et al.· 1 citation· ⚡1
Simulation studies confirm that the federated approach improves estimation and prediction over purely local methods, especially when per-client data are scarce, and an MRI-based ADHD study illustrates its strong performance under real privacy constraints.