A joint model update scheduling and resource allocation framework, aiming to maximize long-term semantic fidelity and resource efficiency of UAV edge intelligence systems, and a Lyapunov-guided discrete reinforcement learning algorithm that performs action space dimensionality reduction and transforms constraints into virtual queue stability problems.
Abstract
While deploying hierarchical vision models to process mission-critical tasks, UAV edge systems must adaptively update the models to sustain inference reliability under low-level environmental corruption. However, existing work has overlooked the optimal timing for model updates, the impracticality of relying on real-time expert labels, and the significant bandwidth and energy constraints of UAVs. This paper proposes a joint model update scheduling and resource allocation framework, aiming to maximize long-term semantic fidelity and resource efficiency of UAV edge intelligence systems. To address the challenge of label-free semantic evaluation, we formulate the Online Semantic Disagreement Rate (OSDR) as a proxy for timely update triggering, thereby enabling fine-grained Sensitivity-Aware Structural Synchronization (SASS). Furthermore, to overcome the curse of dimensionality in hybrid action spaces and effectively bound long-term energy budgets, we propose a Lyapunov-guided discrete reinforcement learning algorithm that performs action space dimensionality reduction and transforms constraints into virtual queue stability problems. The reported experimental results, based on real traffic traces, demonstrate that the proposed framework consistently outperforms representative baselines in semantic recovery efficiency and update triggering precision, by satisfying long-term energy budget and by reducing average risk backlog by up to 33.3\% in the dynamic environmental corruption scenario.
The evolution of uncrewed aerial vehicles (UAVs) into embodied intelligent agents in the low-altitude economy is reshaping edge computing networks. However, the high mobility of UAVs induces severe topology dynamics, limiting the efficacy of traditional fully connected multiagent reinforcement learning because of dimensionality and credit assignment challenges. Furthermore, existing graph attention approaches neglect explicit communication boundary constraints, leading to mismatches between value evaluation and physical topology. To address these challenges, this paper proposes the spatial-aware graph attention multiagent twin delayed deep deterministic policy gradient (SAGA-MATD3) algorithm. By embedding a dynamic spatial masking mechanism based on the communication radius into the critic network, the proposed method enforces physical reachability constraints and attenuates extraneous noise. Simulation results demonstrate that SAGA-MATD3 significantly reduces service latency and improves fairness under an acceptable energy-consumption tradeoff, achieving a 37.1% improvement in convergence reward and enabling the self-organization of robust load-balanced mesh topologies.
Ye Wang, Jingjing Wang, Jianrui Chen et al.· IEEE Transactions on Cogniti...· 0 citations
Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecutive layer groups to different nodes, but the large volume of intermediate activation data incurs substantial transfer overhead that can erase its benefit. Token compression reduces downstream computation and activation transfer, but its quality impact depends on input content, model depth, and earlier pruning decisions, while layer offloading must adapt to time-varying contact and battery conditions. We present \sys, a content-aware hierarchical scheduler that screens constellation-wide options to retain a bounded candidate set, then refines each candidate into a complete token compression and layer offloading trajectory using quality prediction and joint planning. A unified objective balances per-task latency, energy, and quality loss against accumulated workload and battery pressures. We implement \sys on an NVIDIA Jetson AGX Orin hardware-in-the-loop testbed and use its validated execution model for constellation-scale trace replay across multiple ViT workloads and constellation settings. At \(5\)~tasks/s, \sys accomplishes 91.6\% of released tasks, 26.1 percentage points above MARATD3, the strongest baseline, while reducing mean latency and battery draw by 53.0\% and 70.8\%, respectively, and meeting quality targets.
Yan Chen, Yun-Xiang Zhang, Guan-Jun Jiang et al.· 0 citations
: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.
Fei-Yan Bu, Zheng Wang, Yong Pan et al.· Computers, Materials & C...· 0 citations
A UAV motion-aware image capturing and communication system that dynamically optimizes data offloading by jointly considering scenario variability and communication resource allocation and a fast learning-based optimization algorithm (FLO-MICC).
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
6G mobile edge networks are emerging as a key infrastructure for ubiquitous large language model (LLM) inference services. However, conventional edge routing to nearby or well-connected servers falls short for efficient edge LLM inference, as it may miss the user’s KV cache and trigger costly prefill recomputation. To address this challenge, this paper studies an edge inference system assisted by an embodied UAV agent swarm, where UAVs actively sense user mobility and neighboring UAV states to make local decisions on trajectory control, user association, and inference-request routing. The goal is to improve KV-cache reuse while maintaining reliable wireless connectivity, thereby maximizing the system effective token throughput under energy and QoS constraints. We then formulate the joint optimization as a mixed-integer non-linear program and further cast the sequential UAV decision-making process as a decentralized partially observable Markov decision process. To obtain scalable decentralized policies under partial observations, we propose Q-MAA2C, a quantum-enhanced multi-agent advantage actor-critic algorithm for embodied UAV swarm control and inference routing. Q-MAA2C uses quantum actors for local action selection and an entangled split critic for swarm-level value estimation, enabling coordinated policies from partial observations with reduced raw observation exchange. Simulation results indicate that Q-MAA2C yields comparable reinforcement learning rewards to the fully classical baseline while reducing the number of convergence episodes by about 43%. Additionally, the proposed method enhances the system effective token throughput by up to about 134% over other competing methods.
Xiangdong Zheng, Long Luo, Hongfang Yu et al.· IEEE Transactions on Cogniti...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.