Skip to content
Open access

ProfEdge: Efficient Construction of DNN Performance Evaluation Model on Edge Devices

Sep 2026 · ACM Transactions on Internet Technology · 0 citations · 23 references

TL;DR

Experiments on various edge devices and DNN models, including CNN-based and transformer-based workloads, show that ProfEdge reduces profiling errors by up to 80% and saves over 70% of profiling construction cost compared with existing methods.

Abstract

With the rapid development of the Internet of Things (IoT) and artificial intelligence (AI) technologies, edge computing emerges as a crucial computing paradigm. By processing data near the source, edge computing enables faster and more efficient intelligent services. However, edge devices usually limit computational resources, and existing DNN inference latency profiling methods often rely on internal model details or large-scale latency measurements, making them costly and unsuitable for black-box deployment scenarios. This paper proposes ProfEdge, a fast construction framework for deep neural network (DNN) latency profiling models based on Gaussian process regression and Bayesian optimization. ProfEdge adaptively samples real latency measurements under different batch-size states and tunes profiling-model hyperparameters to reduce construction cost while improving profiling accuracy. Specifically, ProfEdge builds an adaptive sampling module based on Gaussian process regression to locate high-error regions through coarse-grained sampling and dynamically refine the sampling process. It further designs a dynamic Bayesian optimization mechanism to improve the accuracy of the latency profiling model. Finally, ProfEdge constructs a cross-device performance mapping model to migrate an existing profiling model to a target device with lightweight stratified calibration, thereby avoiding full reconstruction of the target-device profiling model. Experiments on various edge devices and DNN models, including CNN-based and transformer-based workloads, show that ProfEdge reduces profiling errors by up to 80% and saves over 70% of profiling construction cost compared with existing methods. The cross-device migration results further demonstrate that ProfEdge can achieve competitive profiling accuracy on new devices with only a small number of target-device calibration samples.

Read PDF

Similar papers

#edge computing Open access Aug 2026

DAPart: An Online DRL-based Adaptive Partition Framework for DNN Inference Acceleration and Energy Conservation in Edge Computing

An online Deep Reinforcement Learning (DRL) based adaptive partition method to dynamically determine optimal partitioning decision so as to jointly accelerate DNN inference and mitigate energy consumption is developed.

Shu-Bin Zhang, Junrong Ma, Kai-Kai Chi et al. · 0 citations

Lightweight Deep Learning Models: Technologies, Applications, and Edge Deployment

This paper systematically sorts out the technical connotation and typical cases of lightweight models, and analyzes their application practices in the fields of medical image analysis, financial real-time risk control, and predictive maintenance of industrial Internet of Things.

Z. Su · 0 citations
Conference Aug 2026

A Lightweight Cross-Feature Slicing Model for Real-Time Resource Allocation in 6G Edge Networks

The 6G wireless networks emerge, the need for realtime, efficient resource allocation in Internet-of-Vehicles (IoV) systems becomes critical. Existing machine learning approaches achieve high accuracy, but produce models of several hundred kilobytes, making them unsuitable for deployment on IoT-class Multi-access Edge...

Paa Jayasinghe, M. Maduranga, Sabyasachi Bhattacharyya et al. · 0 citations
Open access 2026

DVFS-Aware Energy-Efficient Implementation of DNNs on Heterogeneous Hybrid Parallel Computers

High-performance heterogeneous computing has emerged as a key resource for training increasingly Deep Neural Networks (DNNs). Achieving optimal performance and energy efficiency on such platforms, however, requires careful resource optimisation at both application and platform levels. While state-of-the-art DNN optimis...

Hamidreza Khaleghzadeh, Atefeh Khazaei, Alexey L. Lastovetsky · 0 citations
Open access Aug 2026

LLM-Driven Context-Aware Health Monitoring for Resource-Constrained Edge Devices

An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.

Ioannis Tzitzios, A. Dimara, Georgiana Petridou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.