In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines.
Abstract
This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term QoS with limited inter-controller signaling. Each base station locally executes its actor without centralized training or actor federation, while critic knowledge is exchanged peer-to-peer over an interference-aware graph. It enables this collaboration through wireless-aware event triggering, adaptive layer-wise top-$k$ sparse critic exchange with error feedback, and balanced interference-aware fusion. We establish conditional finite-time stationarity and consensus guarantees for the balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. It achieves the highest held-out throughput, improves user-rate distribution and mean SINR, increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It reduces critic-communication overhead by $76\%$ relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate RAN controllers without centralized trajectory collection or parameter-server aggregation.
Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation, higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines.
Jayesh Kumar Dabi, Priyadarshi Ashok Dahat· International Journal of Wir...· 0 citations
Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.
Jinho Choi· IEEE Transactions on Communi...· 0 citations
Uncrewed aerial vehicle (UAV)-enabled federated learning (FL) can provide flexible, on-demand edge intelligence for large-scale IoT deployments, but operating in shared unlicensed bands makes uplink update delivery interference-coupled and unreliable. In this paper, we develop a packet-level transmission framework that captures buffer overflow, delay violations, and transmission errors, and uses the resulting packet delivery ratio (PDR) to represent partial-update reception through a packetized, Bernoulli-masked FL aggregation process. We then formulate a fairness-consensus bilevel (FCB) optimization that jointly controls (i) transmission thresholds to maximize the average PDR while reaching consensus under partial observability and (ii) transmission powers to improve the worst PDR and enforce fairness across IoT learners. To solve this problem, we propose an alternating FCB optimizer composed of a consensus-based threshold controller (CTC), which drives the IoT learners toward a PDR-efficient consensus on transmission thresholds, and a fairness-based power controller (FPC), which updates transmission powers to improve the worst PDR and ensure fairness under the resulting consensus thresholds. Numerical results on CNN-based FL tasks show that the FCB optimizer improves FL aggregation and training performance by enhancing packet-level update delivery, consistently outperforming baseline transmission policies.
Over-the-Air Computation (AirComp) Federated Learning (FL) is actively studied as a communication-efficient technique for distributed Artificial Intelligence (AI) model training. To mitigate the impact of wireless channels on the aggregated global model while addressing client energy sustainability, recent efforts have explored integrating Reconfigurable Intelligent Surfaces (RIS) and Simultaneous Wireless Information and Power Transfer (SWIPT) into AirComp FL. In this context, literature has mainly focused on radio resource allocation for optimized SWIPT and RIS-assisted Downlink (DL) model broadcasting and Uplink (UL) AirComp model aggregation. Nevertheless, existing works largely treat the communication design of AirComp FL in isolation, neglecting the tight coupling between radio and compute resource allocation. In this paper, we address this gap by modeling the radio-compute dependency in AirComp FL and optimizing harvested energy to sustain client-side local training and model transmissions. To this end, we jointly optimize the RIS configuration, SWIPT power-splitting ratio, DL transmission time, and local computing frequency to minimize the total communication and computation overhead in latency and energy. The original non-convex problem is decomposed into two independent subproblems, which are solved iteratively via a combination of low-rank optimization and min-max convex reformulation techniques. Numerical evaluations confirm that integrating RIS and SWIPT into AirComp FL leads to higher accuracy, and reduced latency and energy overheads across the FL pipeline.
Stefanos Voikos, P. Charatsaris, Maria Diamanti et al.· IEEE Transactions on Wireles...· 0 citations
The proposed framework does not optimize only computational speed, but also clarifies the trade-off among execution time, SINR, spectral efficiency, and fairness under dynamic uplink CF-mMIMO conditions, indicating that this architecture serves as an adaptable platform to evaluate dynamic uplink power distribution across CF-mMIMO networks.
Hussein A. Jasim, M. F. A. Rasid, F. Hashim et al.· Engineer· 0 citations
A Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs is proposed, allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements.
Ishan Budhiraja, Abhay Bansal, B. Unhelkar et al.· Transactions on Emerging Tel...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.