2026· IEEE Transactions on Communications· Vol 74, pp. 12183-12196· 0 citations· 44 references
Computer Science
Abstract
Low Earth orbit (LEO) satellite-terrestrial communication systems grapple with significant challenges posed by their inherent dynamism and substantial transmission delays. To address these critical issues, this paper proposes a novel hybrid-medium transmission optimization framework that leverages high-altitude platforms (HAPs) as relays. Our primary objective is to minimize end-to-end system delay through the joint optimization of transmission mode selection and wireless communication resource allocation. The resulting joint optimization problem is formulated as a computationally intractable mixed-integer nonlinear programming (MINLP). We present a hierarchical solution strategy to tackle this complexity. Firstly, Lagrangian optimization is employed to analytically derive the intrinsic coupling between resource allocation and transmission mode selection, thereby simplifying the problem into a sequential decision-making process. This sequential problem is subsequently framed as a Markov decision process (MDP), enabling the design of a deep reinforcement learning (DRL) agent tasked with dynamically learning the optimal transmission mode selection policy. By maximizing cumulative long-term rewards, our DRL-based approach effectively reduces overall system delay, unlocking enhanced performance potential for future 6G networks.
The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
Yassine Afif, Ashutosh Balakrishnan, Philippe Martins et al.· 0 citations
This paper considers a downlink communication framework comprising a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-aided by orthogonal time frequency space (OTFS) and non-orthogonal multiple access (NOMA) technologies. Further, delay-Doppler mobility in such frameworks renders classical alternating optimization impractical for per-coherence interval reconfiguration. To mitigate such issues, the STAR-RIS phase-shift and energy-splitting design is formulated as a constrained, non-convex sum-rate maximization problem with closed-form maximum ratio transmission beamforming and fixed NOMA power allocation. To circumvent the per-interval re-optimization burden, a deep reinforcement learning (DRL) approach is adopted that maps observed channel realizations to STAR-RIS configurations through a single forward pass. Specifically, Beta-Space Soft Actor-Critic (SAC-BSE), a maximum entropy DRL agent, is proposed. Simulation results, with two NOMA-multiplexed users on each STAR-RIS branch, confirm rapid convergence, limit the sum-rate degradation to roughly 10\% across a 128-fold user-speed range, and yield consistent gains over OTFS-only, NOMA-only, STAR-RIS-only, fixed-split, and mode-switching baselines as transmit power and the number of STAR-RIS elements increase.
Rais J. Gachaba, Manobendu Sarker, Anirban Bhowal· 0 citations
This paper studies the fast and high-performance FA reconfiguration for low-altitude FA networks with multi-agent reinforcement learning (MARL) and presents an electromagnetic digital twin (EM-DT)-assisted MARL framework to fill the sim-to-real gap.
Tong Zhang, Yanan Su, Shuai Wang et al.· 0 citations
Designing Low Earth Orbit (LEO) constellations for applications like Positioning, Navigation, and Timing (PNT) is a challenging multi-objective optimization challenge. Conventional metaheuristics often suffer from premature convergence due to their reliance on static adaptive rules, limiting their effectiveness in complex search space. To address this limitation, we proposed a hybrid framework where a Double Deep Q-learning Network (DDQN) agent learns a policy to adaptively control the key parameters of a Particle Swarm Optimization (PSO) algorithm. The proposed framework formulates Walker constellation optimization as an sequential parameter control problem. Based on constellation performance feedback, the DDQN controller jointly selects the inertia weight and acceleration coefficients of PSO, guiding the PSO to more effectively balance exploitation and exploration. In a regional design case for China, our algorithm demonstrated superior performance. Compared to a 120 satellites benchmark constellation, the optimized constellation achieved a 27% reduction in Geometric Dilution of Precision (GDOP), a 27.6% enhancement in navigation accuracy, and a 5% increase in coverage multiplicity. This work establishes a robust methodology for the automated and intelligent design of LEO systems, validating the potential of deep reinforcement learning methods for complex aerospace optimization problems.
Zi-Xuan Rui, Fang-Ling Zeng, Xiao-Feng Ouyang et al.· Italian National Conference...· 0 citations
Artificial intelligence-driven intelligent algorithms have demonstrated excellent adaptive optimization capabilities in complex network scheduling problems. To address the issues of dynamic topological changes and insufficient resource allocation efficiency in LEO satellite inter-satellite communication link scheduling, this study proposes a link scheduling model based on deep reinforcement learning. Building upon a dynamic time-varying network model, a Markov decision process is introduced to describe the scheduling process, and a deep neural network is employed to achieve a nonlinear mapping from states to actions. By integrating delay, throughput, and load balancing metrics through a multi-objective reward function, the scheduling strategy is optimized. Experimental results demonstrate that this method exhibits superior performance in reducing transmission delay, enhancing system throughput, and improving load balancing, thereby validating its effectiveness and adaptability in complex inter-satellite link scheduling scenarios.
Qin Li· Digital Signal and Computer...· 0 citations
Integrated Sensing and Communication (ISAC) is emerging as a key technology for next-generation wireless networks, enabling simultaneous communication and sensing functionalities. This paper focuses a RIS-assisted full-duplex (FD) ISAC system, in which a multi-antenna base station (BS) concurrently performs multi-user uplink and downlink transmission while also carrying out radar sensing. To maximize the joint uplink–downlink sum rate, an optimization problem is formulated under practical constraints, such as radar detection SINR, self-interference, BS transmit power, user power budgets, and RIS unit-modulus conditions. To address the nonconvexity of this problem, a two-stage hybrid optimization approach is developed. In the first stage, the augmented Lagrangian technique decomposes the complex problem into simpler subproblems involving beamforming, power allocation, and RIS phase optimization, leading to a feasible initial solution. The second stage employs a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) framework to refine this solution adaptively, enabling the system to respond effectively to variations in the channel environment, mobility patterns, and interference levels. The proposed hybrid framework achieves optimal resource allocation while maintaining feasibility, robustness, and adaptability. Analytical results confirm its convergence behavior, and extensive simulation results confirm that the proposed scheme consistently outperforms conventional optimization and single-agent DRL baselines in sum-rate maximization, interference mitigation, and sensing accuracy, confirming its effectiveness for RIS-assisted full-duplex ISAC systems.
S. Waqas, Fenghua Huang, Fakhar Abbas et al.· IEEE Transactions on Wireles...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.