Jul 2026· International Conference on Ubiquitous and Future Networks· pp. 567-569· 0 citations· 15 references
Abstract
To support heterogeneous services in 5G and beyond, multi-radio access technology (multi-RAT) mobile edge computing (MEC) systems must jointly handle user association and radio resource allocation under stringent delay requirements. This problem becomes particularly challenging when ultra-reliable low-latency communication (URLLC) and enhanced mobile broadband (eMBB) users coexist, since the former requires strict deadline satisfaction while the latter emphasizes low average latency under computation-intensive workloads. In this paper, we propose a hierarchical multi-RAT MEC framework with large language model (LLM)-assisted evolutionary computation (EC). The outer layer uses an LLM to generate user–base station association decisions, while the inner layer employs EC for bandwidth allocation under the given association pattern. This design combines the global reasoning capability of LLMs with the continuous optimization strength of EC. Simulation results show that the proposed LLM-hRAT scheme converges faster and achieves lower average cost than conventional single-layer EC methods and heuristic non-LLM baselines in both uniform and non-uniform user deployment scenarios.
The next generation of wireless systems extends ultra-reliable low-latency communications (URLLC) to the realm of massive connections, termed mURLLC. To address the inherent conflict between stringent quality of service (QoS) requirements in URLLC and the problem of severe and highly fluctuating interference behind demands of massive connectivity, effective fast fading (FF) mitigation and resource allocation strategies are crucial. Through in-depth analysis of FF characteristics, this paper derives optimized configurations for two FF mitigation approaches: protection margin reservation and $K$ -repetition. Furthermore, we integrate these FF mitigation strategies into a hierarchical-clustering (HC)-based resource allocation algorithm for configured-grant in mURLLC. This results in a highly practical and efficient algorithm for managing radio resources and interference in mURLLC scenarios. Simulation results demonstrate that our proposed algorithm achieves over 65% reduction in resource consumption without compromising reliability, significantly enhancing network capacity to support demanding mURLLC applications.
Yi-Chen Guo, Lili Xu, Yihang Cheng et al.· IEEE Transactions on Wireles...· 0 citations
The coexistence of Ultra-Reliable Low-Latency Communication (URLLC) and Enhanced Mobile Broadband (eMBB) can simultaneously meet the reliability and real-time performance requirements of critical services as well as the requirements of high-bandwidth services. In 5G and beyond 5G (B5G) networks, radio access network (RAN) slicing can ensure the differentiated quality of service (QoS) for coexisting heterogeneous services. This paper investigates a dynamic resource allocation framework, aiming to maximize resource utilization subject to QoS constraints. The optimization problem is an NP-hard integer program. To address this issue, we propose a hierarchical proximal policy optimization (PPO) model assisted by a traffic prediction algorithm, namely graph-aggregated Extended Long Short-Term Memory (GxLSTM), which decouples the resource allocation behavior between eMBB and URLLC. Transfer learning is introduced to improve the learning efficiency, forming a Transfer Learning-assisted Hierarchical PPO algorithm (TL-HPPO). In addition, considering the lightweight algorithm deployment requirements for edge networks, based on angular knowledge distillation (AKD), GxLSTM and TL-HPPO are respectively distilled to obtain AKD-traffic prediction (AKD-TP) and AKD-resource allocation (AKD-RA), thereby reducing computational complexity. Simulation results show that the proposed algorithms outperform comparative algorithms while ensuring QoS, effectively improving resource utilization.
Yixuan Bai, Heng Wang· IEEE Transactions on Communi...· 0 citations
Large language models (LLMs) are increasingly deployed on edge nodes to support edge intelligence applications. To overcome limited GPU memory, offloading-based methods partition model parameters between the GPU and host memory, enabling inference on commodity hardware. However, deploying a single model instance using the offloading-based method often results in significant infrastructure overhead and the underutilization of CPU, GPU, and PCIe resources due to a persistently idle CPU, bursty workload patterns, and bandwidth–compute mismatches. To address this issue, this article proposes RACS, a resource-aware cooperative scheduling (RACS) framework that enables a single edge node to coserve a latency-critical high-priority model and a latency-tolerant low-priority model. The key insight is that PCIe bandwidth constitutes the primary bottleneck in offloading-based inference. RACS comprises a runtime state manager that monitors PCIe availability in real time and a resource-aware cooperative scheduler that orchestrates the low-priority model accordingly. When the high-priority model is active, RACS restricts low-priority execution to preloaded feed-forward layers to avoid PCIe contention. When PCIe is idle, RACS aggressively utilizes GPU and PCIe resources while cooperatively scheduling computations on the CPU to maximize throughput. Extensive experiments with the OPT-13-B and OPT-6.7-B models under diverse prompt lengths, generation lengths, and real-world request traces demonstrate that RACS improves the throughput of offline tasks by up to 27.4% without compromising the latency of the high-priority model.
Zhen-Zheng Li, Zhiqing Tang, Jian-Xiong Guo et al.· IEEE Internet of Things Jour...· 0 citations
This study jointly optimizes task offloading and system resource scheduling to minimize the long-term delay–energy cost of NOMA-MEC systems using a master-refined multi-agent proximal policy optimization algorithm.
An adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC, which uses a Beta policy to model bounded actions and achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings.
Zheng Yao, Jie Liu, Changjun Deng et al.· Computers, Materials & C...· 0 citations
Results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.
Sabeur Lajili, Zaki Brahmi· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.