Jul 2026· International Conference on Edge Computing [Services Society]· pp. 1-9· 0 citations· 46 references
Abstract
Splitting complex model inference between multiple computing devices can overcome latency and energy constraints at the edge. Newer edge accelerator devices with higher computing capacity and energy efficiency, enable more fine-grained offload throughout layers of the network, leading to the potential for multiple split configurations. However, optimizing a DNN for inference across networked devices requires a precise performance model that can guide design choices. In this paper, we demonstrate that existing models for computation and communication latency are inaccurate due to system considerations and propose a new empirical model based on structured benchmarking, considering data ingestion overhead due to transfers between devices as well as within a device for data to reach the GPU. We validate our split inference performance model using VGG16 and ResNet50 networks on two heterogeneous platforms, showing it achieves a mean absolute error of no more than 4% for both DNNs, significantly outperforming previous models with errors of more than 15%. We also validate the practical utility of our model by incorporating it into existing split inference search algorithms under multi-split, dynamic bandwidth, and multi-tenant scenarios, demonstrating its effectiveness in navigating the split inference search space.
This work develops time and energy roofline characterizations for Orin AGX across power modes, coupled with analytical models of compute and memory for DNN inference, and extends these models to DNN training, and demonstrates power-mode tuning that achieves up to \(15\%\) lower energy with small inference-time impact f...
P. K, Kunal Kumar Sahoo, Amartya Ranjan Saikia et al.· Proceedings of the Internati...· 0 citations
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment n...
Maysam Khatib, Moysis Symeonides, Demetris Trihinas et al.· 0 citations
HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package, and describes collective designs that improve compute-communication pipelining for latency-sensitive workloads are presented.
W. Bland, Tiago Antunes, Lars Paul Huse et al.· 0 citations
Deep neural networks (DNNs) with billions of parameters power many important applications, but their training is fundamentally constrained by the limited on-chip memory of GPUs. This memory wall forces training to rely on distributed execution or memory offloading, both of which introduce substantial inefficiencies. Ex...
Xiaoyang Sun, Jie Xu, Zheng Wang· IEEE Transactions on Paralle...· 0 citations
A time-varying integer program to minimize the long-term total cost of the edge AI inference system, including the inference latency, the inference error rate, the query-dispatching communication cost, and the energy consumption, subject to resource and workload constraints is proposed.
Ming-Tao Ji, Hehan Zhao, Lei Jiao et al.· Science China Information Sc...· 0 citations
This work analyzes the Vitis AI compiler and proposes an XIR-level splitting framework that generates independently compilable .xmodel fragments while preserving the context required for DPU mapping, and restores correct DPU mapping by addressing boundary-context loss and incomplete dependency collection.
Federico Buccellato, Luca Mannini, C. De Sio· WiPiEC Journal - Works in Pr...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.