Sep 2026· Turkish Journal of Electrical Engineering and Computer Sciences· 0 citations· 11 references
TL;DR
A reparameterized transformer framework that integrates High-Rank Factorization (HRF) during training, layer merging at inference, and dynamic, load-balanced distributed inference across multiple devices is proposed, highlighting that reparameterized transformers, coupled with adaptive distributed inference and ultra-low-precision execution, offer a compelling solution for real-time, on-device analytics in IoT and edge computing environments.
Abstract
Deploying advanced transformer-based models on resource-constrained edge devices remains a significant challenge due to their high memory footprint and substantial compute requirements. In this paper, we propose a reparameterized transformer framework that integrates High-Rank Factorization (HRF) during training, layer merging at inference, and dynamic, load-balanced distributed inference across multiple devices. To further reduce resource usage, our framework supports mixed-precision quantization down to 4-bit, enabling flexible accuracy–latency–energy trade-offs. Experimental evaluations on the ESC-50 environmental sound dataset demonstrate that our method matches or exceeds the performance of larger baseline models while using 20–30% fewer parameters, achieving up to 48% latency reduction in multidevice setups, and substantially lowering energy consumption. Ablation studies confirm the benefits of tuning factorization rank, partition strategy, and quantization level for diverse edge scenarios. Overall, our findings highlight that reparameterized transformers, coupled with adaptive distributed inference and ultra-low-precision execution, offer a compelling solution for real-time, on-device analytics in IoT and edge computing environments.
This work presents a distributed inference framework that integrates speculative decoding across edge and cloud, and shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement.
D. J. Bajpai, K. Upadhyay, M. Hanawal· 0 citations
DABO is proposed, a calibration-aware binary offloading method for collaborative large–small model inference that maintains competitive end-to-end accuracy while processing an average of 83.72% of requests at the edge.
Chen Zhu, Yi-Ming Su, Chenwenjie Mao et al.· IEEE Access· 0 citations
Collaborative fine-tuning on edge devices adapts large language models to domain-specific data while keeping each device's data local. State-of-the-art (SOTA) collaborative fine-tuning techniques are largely designed for GPU-based edge devices and rely on pipeline parallelism (PP). However, many edge platforms, includi...
Wonmi Choi, Sunjae Park, Dohyeok Kwon et al.· 0 citations
SplitLite is proposed, a communication-efficient split federated LoRA fine-tuning method that exploits the low effective rank structure of consecutive-epoch activation and gradient residuals, thereby significantly reducing both activation uplink and gradient downlink traffic.
Tao Li, Yu-Lin Tang, Qi Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.