Aug 2026· International Journal of Interactive Mobile Technologies (ijim)· 0 citations· 5 references
TL;DR
This study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents and shows decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.
Abstract
Large language models (LLMs) are becoming more and more important for perception, reasoning, navigation, and human-machine interaction in autonomous mobile agents like delivery robots, drones, and intelligent cars. However, the rigorous real-time latency requirements, energy limitations, and restricted processing resources make it difficult to implement LLMs directly on edge devices. Therefore, this study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents. To obtain effective on-device intelligence while preserving effective performance, the suggested method combines lightweight model architectures, adaptive inference scheduling, dynamic job offloading, and hardware-aware optimization techniques. While edge–cloud collaboration allows for the selective execution of computationally demanding activities, quantization, pruning, and knowledge distillation are used to lower model complexity and memory consumption. Furthermore, an energy-aware resource management technique constantly modifies processor workloads according to task urgency, network conditions, and battery levels. When compared to traditional cloud-dependent LLM deployments, experimental evaluations on representative mobile robotic platforms show notable improvements in inference latency and battery consumption. The findings show decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.
The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML). HYMELL models LLM execution through a three-level hierarchy: analytical estimation of primitive operations, ML prediction of higher-level components, and an end-to-end model that captures system-level overheads across both prefill and decode phases. The framework supports diverse architectures, including dense and mixture-of-experts (MoE) feed-forward networks (FFNs), as well as multi-head attention (MHA) and grouped-query attention (GQA) mechanisms. Evaluated on an NVIDIA H100 graphics processing unit (GPU), HYMELL achieves high predictive accuracy; notably, for LLaMA 3 8B, it attains less than 5% error for both prefill and decode phases. By predicting execution costs directly from architectural parameters, it enables fast, hardware-free design space exploration and energy-efficient optimization.
Saeid Shokoufa, Mohammad Erfan Sadeghi, M. Kamal et al.· 0 citations
Low-latency and resource-efficient predictive analytics are essential for mobile edge computing applications, including smartphone-based human activity recognition (HAR). This study presents the Adaptive Latency Prediction and Scheduling (ALPS) framework, which aims to minimize inference latency, hardware energy consumption, and computational overhead while maintaining predictive accuracy. The ALPS framework comprises three primary modules: an adaptive principal component analysis (PCA) mechanism for real-time feature reduction, a lightweight ridge classifier optimized with an L regularization loss 2 function for rapid multi-class activity prediction, and a dynamic task scheduler that minimizes a combined latency-energy cost function to allocate processing tasks across mobile, edge, and cloud layers. Evaluation on the high-dimensional UCI-HAR smartphone dataset (10,299 samples, 561 features) demonstrated that the adaptive feature reduction module reduced the feature space to 68 principal components, resulting in an 87.9% reduction in dimensionality while maintaining 95% cumulative explained variance. Relative to conventional standalone models and isolated optimization baselines, ALPS achieved a 32% reduction in average end-to-end latency (120 ms), a 25% decrease in energy consumption (180 J), and a classification accuracy of 96.4% across six physical activities. The primary contribution of this study is the unified integration of adaptive data compression and distributed infrastructure scheduling into a scalable and energy-efficient pipeline for real-time edge intelligence.
M. Nohong, Nora’asikin Abu Bakar, Siti Roshaida Abd Razak et al.· International Journal of Int...· 0 citations
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Zejian Wang, Nan Lin, Zinuo Cai et al.· ACM Transactions on Architec...· 0 citations
A Mobile Reasoning-as-aService (MORES) framework that treats reasoning as a computational service accessible to edge devices over wireless networks, and focuses on implicit reasoning, which achieves an approximately 18% improvement in system throughput over the baseline Soft Actor-Critic (SAC) algorithm.
Guanchen Liu, Hongyang Du, Kaibin Huang· 1 citation
This study investigates energy-efficient distributed machine learning techniques, including federated learning, model compression, adaptive resource management, dynamic task offloading, and communication-efficient optimization, and proposes a distributed learning framework that integrates local model training, adaptive communication scheduling, gradient compression, and workload balancing to minimize energy consumption while maintaining learning accuracy.
Venkatesh Iyer· International Journal of App...· 0 citations
An LLM-driven, context-aware framework that integrates real-time system metrics, historical data, and task-specific importance levels for anomaly detection and prediction is proposed, enabling proactive intervention before critical operating conditions are reached.
Ioannis Tzitzios, A. Dimara, Georgiana Petridou et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.