Skip to content

Hybrid Learning Framework for Throughput Enhancement in IoT Networks With Semi-Grant-Free NOMA

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 10481-10497 · 0 citations · 29 references
Computer Science

Abstract

Semi-grant-free non-orthogonal multiple access (SGF-NOMA) schemes group one grant-based (GB) user with multiple grant-free (GF) users into one time/frequency resource block (RB) to enhance spectral efficiency. Due to the sporadic traffic of GF users and the stringent quality of service (QoS) requirement of the GB user, the access collision problem becomes severe in SGF-NOMA. To solve this problem, this paper firstly designs an RB-based power pool (PP), which directs GF users to adjust their transmit power without disrupting the ongoing transmission of the internal GB user. After that, this work proposes an efficient multi-agent deep reinforcement learning (MA-DRL) framework to jointly optimize the PP and access strategy for maximizing the network throughput. In particular, this work exploits the fast-response feature of the traditional competitive MA-DRL and the increased-performance feature of the traditional cooperative MA-DRL to redesign a mixed reward system, which contributes to a hybrid MA-DRL mode for enhancing the learning efficiency of agents, i.e., GF users. We investigate the performance of the proposed algorithm at the network level and the NOMA-cluster level. We show that the proposed hybrid MA-DRL at the cluster level converges faster to an optimal solution than that at the network level but at an extra cost of user clustering. The numerical results show that the proposed scheme increases the successful decoded users by 42.38% when compared to the traditional schemes without learning capability. The proposed hybrid MA-DRL mode performs better than the pure competitive and cooperative MA-DRL modes, especially under a heavy-load network. It is able to achieve a 69% success rate of access in a time-varying environment with high packet arrival rates.

View source

Similar papers

Open access Aug 2026

JATO: Deep Reinforcement Learning-based Joint Optimization for Task Offloading and Adaptive Transmission in Multimedia IoT Systems

As the Multimedia Internet of Things (M-IoT) evolves, the orchestration of numerous resources that offer support for high-bandwidth, low-latency applications arises as a key challenge. Architecturally, the edge-cloud framework alleviates structural concerns, but the linked nature of compute and data transfer poses problems of resource management. Approaches that tackle task offloading and adaptive transmission that think independently of each other tend to have problems such as user-server cross-region overloads or network congestion. This paper presents JATO, a framework to jointly tackle the problems of adaptive task offloading and transmission optimization using Deep Reinforcement Learning. JATO offers a mono-faceted solution, learning a policy to simultaneously determine the best offloading target and the transmission quality. The framework was implemented for evaluation with a combination of different edge devices in a testbed alongside a simulation environment. JATO recorded a result of 0.9321 as the holistic score of the overall framework endpoint, a score significantly better than that of all the other frameworks that were used as functional baselines. JATO was able to resource optimally with a network lag of 131.65 milliseconds and a network freeze of 0.09% with the resources utilized. This is evidence that offloading and rate control in combination provides better resource elasticity for M-IoT systems.

G. Purnama, Irma Amelia Dewi, A. Langi et al. · 0 citations
Open access Aug 2026

Energy-Efficient Cooperative Data Offloading in Cellular Networks Using Reinforcement Learning

This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.

Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al. · 0 citations
Conference Jul 2026

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

Yu-Heng An · 0 citations
Preprint Sep 2026

Deep Reinforcement Learning for Optimization of STAR-RIS Phase and Energy Splitting Coefficients in OTFS-NOMA Framework

This paper considers a downlink communication framework comprising a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS)-aided by orthogonal time frequency space (OTFS) and non-orthogonal multiple access (NOMA) technologies. Further, delay-Doppler mobility in such frameworks renders classical alternating optimization impractical for per-coherence interval reconfiguration. To mitigate such issues, the STAR-RIS phase-shift and energy-splitting design is formulated as a constrained, non-convex sum-rate maximization problem with closed-form maximum ratio transmission beamforming and fixed NOMA power allocation. To circumvent the per-interval re-optimization burden, a deep reinforcement learning (DRL) approach is adopted that maps observed channel realizations to STAR-RIS configurations through a single forward pass. Specifically, Beta-Space Soft Actor-Critic (SAC-BSE), a maximum entropy DRL agent, is proposed. Simulation results, with two NOMA-multiplexed users on each STAR-RIS branch, confirm rapid convergence, limit the sum-rate degradation to roughly 10\% across a 128-fold user-speed range, and yield consistent gains over OTFS-only, NOMA-only, STAR-RIS-only, fixed-split, and mode-switching baselines as transmit power and the number of STAR-RIS elements increase.

Rais J. Gachaba, Manobendu Sarker, Anirban Bhowal · 0 citations
2026

Structured Reinforcement Learning for User Admission in Multi-Cell Massive MIMO via O-RAN

Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.

Jinho Choi · 0 citations
Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.