Skip to content
Open access

Federated and Communication-Efficient Decentralized Meta-Reinforcement Learning for Dynamic Spectrum Access in Cognitive Radio–Enabled 5G IoT Networks

Aug 2026 · International Journal of Wireless and Microwave Technologies · Vol 16, pp. 275-290 · 0 citations

TL;DR

Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation, higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines.

Abstract

Dynamic spectrum access (DSA) in 5G IoT setups with cognitive radio is characterized by rapid and decentralized decision-making processes in highly non-stationary wireless environments, limited communication needs, and restrictive bounds. In this work, we present F-DMRL, a federated, communication-efficient decentralized meta-reinforcement learning framework for allowing a massive number of IoT devices to meta-learn collectively about spectrum-access strategies in a decentralized way without centralized control and without an extensive amount of inter-agent communication. Our method incorporates lightweight federated meta-parameter aggregation with gradient sparsification and periodic communication, allowing devices to only compress the meta-updates during this process and then adapt locally for task-specificity. We have presented analytical speedup guarantees and upper bounds on communication cost under bounded environmental drift and shown that using the approach proposed here, F-DMRL preserves convergence properties while posing a large reduction in coordination overhead at the same time. Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation (up to 45% fewer episodes), higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines. Simulation results averaged across 10 independent runs demonstrate improvements of 45% faster adaptation and 60–80% lower communication overhead relative to baseline methods, while maintaining stable convergence.

Read PDF

Similar papers

Preprint Aug 2026

Partially-Observable Transmission Control for UAV-Enabled Federated Learning in IoT Networks

Uncrewed aerial vehicle (UAV)-enabled federated learning (FL) can provide flexible, on-demand edge intelligence for large-scale IoT deployments, but operating in shared unlicensed bands makes uplink update delivery interference-coupled and unreliable. In this paper, we develop a packet-level transmission framework that captures buffer overflow, delay violations, and transmission errors, and uses the resulting packet delivery ratio (PDR) to represent partial-update reception through a packetized, Bernoulli-masked FL aggregation process. We then formulate a fairness-consensus bilevel (FCB) optimization that jointly controls (i) transmission thresholds to maximize the average PDR while reaching consensus under partial observability and (ii) transmission powers to improve the worst PDR and enforce fairness across IoT learners. To solve this problem, we propose an alternating FCB optimizer composed of a consensus-based threshold controller (CTC), which drives the IoT learners toward a PDR-efficient consensus on transmission thresholds, and a fairness-based power controller (FPC), which updates transmission powers to improve the worst PDR and ensure fairness under the resulting consensus thresholds. Numerical results on CNN-based FL tasks show that the FCB optimizer improves FL aggregation and training performance by enhancing packet-level update delivery, consistently outperforming baseline transmission policies.

Masoud Ghazikor, Ni-Zhen Zhou, Morteza Hashemi · 0 citations
Review Open access Aug 2026

Stability-Aware and Feasibility-Sensitive Aggregation in Federated Reinforcement Learning for Edge-IoT Systems: A Review

Federated Reinforcement Learning (FRL) provides a useful basis for distributed policy learning in Edge-IoT systems, where clients interact with local environments without transferring raw operational data to a central server. Yet aggregation becomes difficult when clients operate under different transition dynamics, workloads, resource capacities, communication conditions, and operational constraints. In these settings, local policy updates may not differ only in magnitude or direction; they may also differ in stability, reliability, resource support, and operational feasibility. Conventional averaging is therefore limited, since it does not distinguish stable and feasible updates from unstable or constraint-violating ones. This paper examines aggregation stability and feasibility-sensitive aggregation in FRL for Edge-IoT systems. It reviews and synthesizes related literature across four connected streams: heterogeneous Federated Learning, FRL-based edge decision-making, constrained and safe Reinforcement Learning, and adaptive or reliability-aware aggregation. The reviewed studies are analyzed through six dimensions: learning paradigm, type of heterogeneity, role of policy learning, treatment of operational constraints, aggregation strategy, and whether local feasibility signals influence global aggregation weights. The analysis indicates that existing studies provide valuable foundations, but they usually treat heterogeneity, constraint handling, and aggregation adaptation as separate concerns. The paper identifies a need for aggregation mechanisms that jointly account for update stability, update reliability, resource availability, and constraint feasibility. It positions aggregation as adaptive client influence regulation rather than passive averaging in future Edge-IoT FRL systems.

Majid A. Aslan, A. Al-Shalabi, Ahmed S. Alhegami · 0 citations
#reinforcement learning Open access Aug 2026

Asynchronous Federated Reinforcement Learning for Adaptive Resource Slicing and Low-Latency Task Offloading in Heterogeneous 6G Edge Computing Networks

AF-EdgeRL is proposed, a novel Byzantine-resilient Asynchronous Federated Reinforcement Learning framework tailored for distributed resource allocation and dynamic task offloading and establishes theoretical convergence guarantees under non-convex reinforcement learning objectives.

Daniel Merrow, Tember L. Nair, Lucas Farnandez · 0 citations
Aug 2026

Communication-Aware Federated Learning for Energy Management in Edge-Cloud Autonomous Systems

The proposed framework is validated by conducting simulation-based experiments on the benchmark datasets and synthetic autonomous workloads, where the novelty lies in the design of the system-level federated learning architecture, instead of the datasets themselves.

Jyotsnarani Tripathy, D. Rajalakshmi, A. N. Ramya Shree et al. · 0 citations
Preprint Aug 2026

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines.

A. Farajzadeh, Melike Erol-Kantarci · 0 citations
Open access Aug 2026

Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication

Collaborative training in distributed semantic communication (DSC) networks typically relies on decentralized federated learning (DFL). However, pushing topology-agnostic aggregation into heterogeneous, multi-task environments creates a fundamental bottleneck: it drives negative transfer and overconsensus bias (OCB). This paper introduces a personalized DSC framework that cuts off this cross-task interference. At the node level, a policy-driven multi-path routing mechanism separates task-specific features from shared representations to preserve local fidelity. Across the network, we deploy a"communicationwhile- aggregation"protocol. It calibrates a column-stochastic consensus matrix using task affinities. This limits the system to absorbing complementary knowledge while actively blocking mismatched parameter updates. To bound the convergence, we derive a unified Lyapunov drift analysis. We reveal a strict Ushaped trade-off: deeper topological mixing reduces variance but amplifies structural OCB. Resolving this tension yields a closed-form expression for the optimal aggregation depth. We evaluate the proposed framework on NYU-v2, where the results reveal a clear trade-off between insufficient aggregation and excessive topological mixing. At the analytically derived optimal aggregation depth, our method achieves a 4.77% global relative improvement over the no-aggregation baseline and outperforms decentralized FedAvg, FedAMP, and heuristic max aggregation. We further evaluate the framework on Taskonomy and imperfect wireless links to examine the effects of network-size variation and wireless-link reliability.

Linqi Yin, Tie-Jun Lv, Weicai Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.