Skip to content
Preprint

FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines.

Abstract

This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term QoS with limited inter-controller signaling. Each base station locally executes its actor without centralized training or actor federation, while critic knowledge is exchanged peer-to-peer over an interference-aware graph. It enables this collaboration through wireless-aware event triggering, adaptive layer-wise top-$k$ sparse critic exchange with error feedback, and balanced interference-aware fusion. We establish conditional finite-time stationarity and consensus guarantees for the balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. It achieves the highest held-out throughput, improves user-rate distribution and mean SINR, increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It reduces critic-communication overhead by $76\%$ relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate RAN controllers without centralized trajectory collection or parameter-server aggregation.

View source

Similar papers

Open access Aug 2026

Federated and Communication-Efficient Decentralized Meta-Reinforcement Learning for Dynamic Spectrum Access in Cognitive Radio–Enabled 5G IoT Networks

Simulations across various 5G IoT spectrum environments showed that F-DMRL performed faster adaptation, higher spectral efficiency, and lower interference probability compared to centralized meta-RL, federated DRL, and traditional decentralized RL baselines.

Jayesh Kumar Dabi, Priyadarshi Ashok Dahat · 0 citations
2026

Structured Reinforcement Learning for User Admission in Multi-Cell Massive MIMO via O-RAN

Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.

Jinho Choi · 0 citations
Preprint Aug 2026

Partially-Observable Transmission Control for UAV-Enabled Federated Learning in IoT Networks

Uncrewed aerial vehicle (UAV)-enabled federated learning (FL) can provide flexible, on-demand edge intelligence for large-scale IoT deployments, but operating in shared unlicensed bands makes uplink update delivery interference-coupled and unreliable. In this paper, we develop a packet-level transmission framework that captures buffer overflow, delay violations, and transmission errors, and uses the resulting packet delivery ratio (PDR) to represent partial-update reception through a packetized, Bernoulli-masked FL aggregation process. We then formulate a fairness-consensus bilevel (FCB) optimization that jointly controls (i) transmission thresholds to maximize the average PDR while reaching consensus under partial observability and (ii) transmission powers to improve the worst PDR and enforce fairness across IoT learners. To solve this problem, we propose an alternating FCB optimizer composed of a consensus-based threshold controller (CTC), which drives the IoT learners toward a PDR-efficient consensus on transmission thresholds, and a fairness-based power controller (FPC), which updates transmission powers to improve the worst PDR and ensure fairness under the resulting consensus thresholds. Numerical results on CNN-based FL tasks show that the FCB optimizer improves FL aggregation and training performance by enhancing packet-level update delivery, consistently outperforming baseline transmission policies.

Masoud Ghazikor, Ni-Zhen Zhou, Morteza Hashemi · 0 citations
2026

Radio and Compute Resource Allocation for SWIPT and RIS-Assisted AirComp Federated Learning

Over-the-Air Computation (AirComp) Federated Learning (FL) is actively studied as a communication-efficient technique for distributed Artificial Intelligence (AI) model training. To mitigate the impact of wireless channels on the aggregated global model while addressing client energy sustainability, recent efforts have explored integrating Reconfigurable Intelligent Surfaces (RIS) and Simultaneous Wireless Information and Power Transfer (SWIPT) into AirComp FL. In this context, literature has mainly focused on radio resource allocation for optimized SWIPT and RIS-assisted Downlink (DL) model broadcasting and Uplink (UL) AirComp model aggregation. Nevertheless, existing works largely treat the communication design of AirComp FL in isolation, neglecting the tight coupling between radio and compute resource allocation. In this paper, we address this gap by modeling the radio-compute dependency in AirComp FL and optimizing harvested energy to sustain client-side local training and model transmissions. To this end, we jointly optimize the RIS configuration, SWIPT power-splitting ratio, DL transmission time, and local computing frequency to minimize the total communication and computation overhead in latency and energy. The original non-convex problem is decomposed into two independent subproblems, which are solved iteratively via a combination of low-rank optimization and min-max convex reformulation techniques. Numerical evaluations confirm that integrating RIS and SWIPT into AirComp FL leads to higher accuracy, and reduced latency and energy overheads across the FL pipeline.

Stefanos Voikos, P. Charatsaris, Maria Diamanti et al. · 0 citations
Open access Jul 2026

Dynamic Uplink Power Control for Cell-Free Massive MIMO

The proposed framework does not optimize only computational speed, but also clarifies the trade-off among execution time, SINR, spectral efficiency, and fairness under dynamic uplink CF-mMIMO conditions, indicating that this architecture serves as an adaptable platform to evaluate dynamic uplink power distribution across CF-mMIMO networks.

Hussein A. Jasim, M. F. A. Rasid, F. Hashim et al. · 0 citations
Open access Aug 2026

Quantum Federated Reinforcement Learning‐Based Traffic Offloading and Resource Allocation for RSMA‐Enabled Space–Air–Ground Integrated Networks

A Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs is proposed, allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements.

Ishan Budhiraja, Abhay Bansal, B. Unhelkar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.