A Bidirectional-AoI-Aware Multi-Agent Deep Reinforcement Learning Framework for Vehicular Platooning in Segmented Waveguide-Based Pinching Antenna Systems
2026· IEEE Transactions on Cognitive Communications and Networking· Vol 12, pp. 10965-10978· 0 citations· 49 references
Computer Science
Abstract
Ensuring reliable and low-latency vehicle-to-everything (V2X) communications in high-speed transport settings remains a significant challenge due to severe path loss brought about by non-line-of-sight (NLoS) and coverage gaps in conventional cellular infrastructure. While dielectric waveguide-based pinching antenna (PA) systems have been proposed to mitigate these physical limitations, they suffer from substantial in-waveguide attenuation over long distances. To address these challenges, we propose a segmented waveguide-enabled pinching-antenna (SWAN) architecture in platoon-based V2X networks. By employing dynamic segment selection, SWAN maintains robust line-of-sight (LoS) connectivity while mitigating the in-waveguide attenuation inherent in conventional PA structures. We formulate a joint resource allocation (RA) and mode selection problem to minimise the age of information (AoI) for both uplink platoon monitoring and downlink traffic broadcasting, whilst ensuring the exchange of intra-platoon cooperative awareness messages (CAMs) and minimising power consumption. To solve this high-dimensional problem, we propose a decomposed multi-agent deep deterministic policy gradient (DE-MADDPG) algorithm augmented with twin delayed (TD3) critics by considering each vehicle platoon (VP) as an agent. This approach decouples system-wide coordination from local executions of VPs, enabling efficient learning in dynamic environments. Extensive simulations demonstrate that the proposed framework significantly outperforms standard reinforcement learning (RL) baseline methods, achieving near-optimal uplink and downlink AoI performance, with an average gap of 4.8% to exhaustive search, and near-perfect CAM delivery probability (CDP), which approaches 100%, even under dense traffic conditions.
Simulation results show that the proposed DH-MDRL framework outperforms conventional schemes without IRSs and achieves an excellent trade-off between V2V link constraints’ satisfaction probability and V2I link sum data rates compared to centralized resource allocation approaches.
This paper studies the fast and high-performance FA reconfiguration for low-altitude FA networks with multi-agent reinforcement learning (MARL) and presents an electromagnetic digital twin (EM-DT)-assisted MARL framework to fill the sim-to-real gap.
Tong Zhang, Yanan Su, Shuai Wang et al.· 0 citations
A Large Language Model-guided cooperative decision-making framework for joint velocity control and bidirectional channel selection in an AAM system with Aerial Vehicles communicating with ground Base Stations while following predefined linear routes is proposed.
Qingyang Li, Adnan Quadri, Hongxiang Li et al.· IEEE Access· 0 citations
This paper leverages Open RAN to manage V2X communication and proposes a multi-agent reinforcement learning (MARL) resource-aware system that aims to mitigate interference, optimize resource usage, and enhance quality of service by optimally selecting between sidelink and network transmissions.
M. Barbosa, K. Dias· IEEE Transactions on Vehicul...· 0 citations
Urban logistics are increasingly strained by dynamic traffic conditions and complex operational constraints, rendering traditional optimization methods for the multidepot capacitated vehicle routing problem inadequate. This paper introduces TAWJEEH, a novel hybrid framework that integrates deep reinforcement learning, classical heuristics, and cellular vehicle-to-everything (C-V2X) communications for time-dependent MDCVRP optimization. The framework employs a deep Q network to learn adaptive policies for customer-to-vehicle assignment, complemented by clustering algorithms for initial customer grouping and heuristics for route refinement. Leveraging real-time data streams from C-V2X messages, TAWJEEH dynamically adjusts routes in response to live traffic conditions. Extensive and realistic simulations using SUMO on Hamburg and Luxembourg road networks validate our approach. The performance of TAWJEEH is benchmarked against the Clarke-Wright savings (CWS) heuristic and ant colony optimization (ACO). Results show significant and consistent reductions in key performance metrics; for instance, in the large-scale Luxembourg scenario, TAWJEEH reduces total travel distance by up to 55.4% compared with CWS and 8.9% against ACO. These improvements translate to substantial reductions in cumulative travel time, fuel consumption, CO
2
emissions, and overall operational costs. TAWJEEH proves to be a robust, scalable, and computationally efficient solution, highlighting the potential of combining advanced artificial intelligence techniques with vehicular communication technologies to address complex urban logistics challenges.
Yacine Harkat, Mustapha Hemis, El-sedik Lamini et al.· Transportation Research Reco...· 0 citations
During disaster scenarios and periods of extreme data demand in next-generation wireless communications, conventional terrestrial networks (TNs) often become unreliable or fail entirely, leading to critical service disruptions. In such contexts, non-terrestrial networks (NTNs) emerge as a promising solution to ubiquitous and resilient connectivity. Furthermore, the open radio access network (ORAN) paradigm facilitates network disaggregation and flexible functional splitting among its key components, namely the central unit (CU), distributed unit (DU), and radio unit (RU), which can be deployed across heterogeneous NTN platforms according to service requirements. However, this flexibility introduces significant challenges in terms of network complexity and real-time control. To address these challenges, this paper proposes an intelligent ORAN-enabled NTN framework for emergency communication scenarios. The proposed system leverages graph neural networks (GNNs) to model the dynamic network topology and employs a reinforcement learning (RL)-based Q-learning algorithm, formulated as a Markov decision process (MDP), to enable adaptive and real-time network control. In this framework, network nodes are treated as states, and optimal decisions are learned based on system dynamics. The spatial distribution of user equipment (UE) is modeled using an inhomogeneous Poisson point process (IPPP) with a rejection sampling technique, capturing realistic user density variations. Simulation results demonstrate that the proposed GNN-enhanced RL approach significantly improves network performance in terms of latency and service reliability, thereby enabling efficient and robust operation under emergency conditions.
Unknown authors· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.