Skip to content
Conference

Twin-Guided Meta Learning for Generalizable UAV Trajectory Planning in Low-Altitude Wireless Networks

Jul 2026 · International Conference on Computer Communications and Networks · pp. 1-9 · 0 citations · 23 references

Abstract

Ensuring QoS provisioning in low-altitude wireless networks requires UAV positioning and navigation strategies that adapt to dynamic environments and generalizes across heterogeneous network scenarios. This paper proposes a digital twin (DT)-assisted meta reinforcement learning framework for multi-agent UAV trajectory planning. A high-fidelity network DT serves as a supervisory layer to generate key performance indicators (KPIs) and fine-grained channel knowledge, which guides both domain-specific learning and cross-domain validation. Building on the twin-informed UAV landmarks, we then develop a weakness-aware meta learning scheme: in the inner loop, agents are trained cooperatively toward the self-discovered landmarks under dynamic conditions; in the outer loop, navigation policies are evaluated via the DT to identify bottlenecks and generate targeted hard scenarios, enabling robust adaptation across diverse scenarios. Extensive simulations show that our framework achieves up to 4× higher service coverage compared to baselines, while the target-aware outer-loop adaptation further improves cross-scene performance and model generalization.

View source

Similar papers

Conference Jul 2026

Double Deep Reinforcement Learning–Based UAV Positioning for Throughput Optimization in Wireless Networks

This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.

Berke Kilinç, M. Ö. Efe · 0 citations
Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
Review Jul 2026

Cooperative Multi-UAV Navigation in Complex Environments via Systematic Multi-Agent Deep Reinforcement Learning

A multi-agent deep reinforcement learning framework that addresses issues through coordinated exploration, demonstration exploitation, safe curriculum scheduling, and structure-aware generalisation is proposed, demonstrating strong performance in collaboration success rate, navigation robustness, zero-shot cross-scenario generalisation, and dynamic environment adaptability.

Yuhuang Su, Nabil Aouf · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations
Jul 2026

Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach

The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kinematics and multi-tier network handovers. While Deep Reinforcement Learning (DRL) offers rapid tactical control, it lacks the zero-shot strategic reasoning required to quickly adapt to dynamic Integrated Terrestrial and Non-Terrestrial Networks (ITNTNs). Conversely, Large Language Models (LLMs) excel at semantic reasoning but suffer from high inference latency, rendering them unsuitable for real-time aerodynamic control. To bridge this gap, we propose a novel Hierarchical LLM-driven control framework. A massive cloud-based LLM deployed on a High-Altitude Platform Station (HAPS) manages slow-timescale global load balancing, while lightweight edge-LLMs on individual UAVs translate local observations into tactical sub-goals. These sub-goals guide a fast-timescale physical DRL controller to execute collision-free, handover-aware trajectories. Simulation results demonstrate that our agentic architecture significantly reduces collision rates and improves aggregate system throughput compared to existing baselines.

Zijiang Yan, Hao Zhou, W. Jaafar et al. · 0 citations
Conference Open access 2026

Learning-Guided Symbolic Solver Selection for Dynamic Multi-UAV Missions in Simulation

: Multi-unmanned Aerial Vehicle (multi-UAV) missions require task assignment strategies that adapt to dynamic urgency, threats, and resource constraints. Fixed heuristics lack flexibility, while end-to-end learned policies often omit explicit safety and verification mechanisms. This paper proposes a learning-guided neuro-symbolic orchestration framework in which reinforcement learning operates at a meta level to select among heterogeneous symbolic solvers within a closed simulation loop. Candidate assignments are regulated through a dual-layer verification mechanism comprising a feasibility filter and an episode-level simulator-based evaluator. The meta-decision layer is trained using a Double Deep Q-Network (Double DQN) with rewards reflecting completion, expiration, battery consumption, and threat exposure. Experiments across diverse scenarios demonstrate the effectiveness of adaptive solver selection and improved safety–efficiency trade-offs compared to fixed baselines.

Muhyun Byun, S. Doo, Eunae Lee · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.