It is demonstrated that the reinforcement learning-based approaches, namely Q-learning and SARSA (State-Action-Reward-State-Action), consistently outperform random selection in terms of total channel capacity, attacker detection accuracy, and performance stability.
Abstract
In next-generation wireless networks, communication systems are expected to go beyond simple data transmission and simultaneously provide high data rates, efficiency, and security. This requirement has motivated the extensive adoption of machine learning methods to develop intelligent and real-time network management frameworks, enabling the system to continuously monitor and react to channel variations and user behavior while maintaining efficient information delivery. In this context, the integration of machine learning with beamforming enables adaptive and data-driven beam direction selection, improving both the efficiency and security of wireless links. In this work, a 3GPP-based system model is first implemented under a no-attacker scenario, and an exhaustive search is employed as a reference to identify the best beamforming configurations. The proposed framework is then evaluated in the presence of an attacker and under different network scalability conditions. We demonstrate that the reinforcement learning-based approaches, namely Q-learning and SARSA (State-Action-Reward-State-Action), consistently outperform random selection in terms of total channel capacity, attacker detection accuracy, and performance stability. Among the evaluated reinforcement learning methods, Q-learning achieves the best overall trade-off between detection accuracy and computational efficiency. Our results indicate that the proposed framework provides a stable, scalable, and effective solution for joint beamforming and security-aware decision-making in dynamic and adversarial wireless environments.
With the advent of seventh-generation (7G) wireless systems, the spectrum environment is extremely dynamic and heterogeneous, and traditional methods of sensing do not offer reliable and efficient performance. This paper introduces a self-evolving spectrum sensing system to enable adaptive and intelligent spectrum access based on a deep reinforcement-based learning paradigm. The framework depicts the sensing process as a sequence decision problem where an autonomous agent is able to continually refine its policy as it engages with the environment. The multi-objective reward formulation is designed to maximize the combination of the detection accuracy, false alarms, energy usage, and using the spectrum. Moreover, an adaptive representation of state mechanism is also introduced to indicate the temporal changes and short-term changes in the spectrum occupancy. The self-evolution strategy proposed adjusts learning parameters and decision policies in a dynamic manner that gives a robust operating in a non-stationary environment. The overall analysis of the experiment shows that the structure achieves detection probability of 97.1, false alarm rate reduced to minimum of 3.8, spectral usage maximized to above 92 and convergence rate is quicker as compared to the existing techniques. These results confirm the appropriateness of the proposed method in overcoming the issues of the next-generation wireless systems.
A.L Sriram, H. N. Divya, S. Sabarinathan et al.· International Conference on...· 0 citations
Motivated by the growing vulnerability of integrated sensing and communication (ISAC) systems to jamming attacks, this work explores multistatic architectures as a promising approach to improve robustness in adversarial environments, specifically in the sensing functionality. We consider a multistatic ISAC framework in which a transmit base station (BS) simultaneously performs communication and sensing, while a set of spatially distributed receiver BSs cooperatively process the reflected echoes from the target through a fusion center under sensing-directed jamming. To fully leverage the spatial diversity in this architecture, we formulate a joint optimization problem that maximizes the fused echo signal-to-interference-plus-noise ratio (SINR), subject to the BS power budget and per-user quality-of-service (QoS) constraints by determining the precoding vector at the transmit BS and the receive beamforming vector at the receiving BSs. The resulting problem is non-convex due to the complex coupling between variables, making it challenging to solve using traditional optimization methods. To address this, we propose a gradient-based meta-learning (GML) approach that learns efficient update rules for rapid convergence. Simulation results validate the effectiveness of the proposed approach, achieving up to 95.68 percent of the optimal performance while substantially mitigating the impact of jamming. The findings highlight that multistatic reception with meta-learning offers a scalable solution to security threats in ISAC systems.
Aya El Bokhary, Ali Amhaz, S. Sharafeddine et al.· IEEE Wireless Communications...· 0 citations
Deep learning is a promising approach to optimize wireless communication by simplifying the search for near-optimal solutions. Prior studies on deep learning-based wireless communication optimization have explored supervised learning approaches that map raw user information, such as location or channel state information, to optimal power allocation vectors. While this approach demonstrates competitive performance, it is susceptible to adversarial attacks via input perturbations. Current defense mechanisms primarily rely on empirical methods, which do not provide formal guarantees of robustness. We fill this gap by proposing a formal verification framework to evaluate the robustness of deep learning-based power allocation in multi-cell massive multiple-input multiple-output (MIMO) systems against a wide range of potential adversarial input manipulations. To the best of our knowledge, this is the first attempt to formally verify deep neural networks in a regression setting with non-linear output constraints. We model the adversary's capabilities using hyper-rectangle constraints on their perturbation, adopt the abstraction-based bound-propagation technique (DeepPoly) to bound the interval of potential allocated powers, and formulate the minimum performance requirements as a constrained program for numerical feasibility analysis. Evaluation on publicly available datasets for power allocation in multi-cell massive MIMO indicates that a well-trained model can guarantee the local robustness under location perturbation by +-1m while retaining a maximum 1% optimality gap.
Thanh Le, Takeshi Matsumura, Yusheng Ji et al.· arXiv.org· 0 citations
Cooperative Spectrum Sensing (CSS) is a promising approach allowing secondary users to utilize unused spectrum without interfering with primary users. However, the collaborative nature of CSS makes it vulnerable to malicious nodes that inject falsified sensing data. Existing attacks have incorporated AI to enhance the effectiveness and overcome challenges in opaque-box settings. However, conventional machine learning models often fail to adapt to rapidly changing wireless environments. In this paper, we propose a large language model (LLM)-powered multi-agent attack framework that leverages the reasoning and adaptability of LLMs to coordinate multiple agents for generating adaptive and context-aware fake sensing reports. The proposed architecture consists of a generator and a discriminator, which work collaboratively to progressively refine falsified sensing data. In addition, to improve prompting efficiency, we incorporate wireless-domain knowledge into the generator to enable advanced prompting, thereby ensuring more effective and efficient malicious report generation. We implement the proposed attack framework in a TV white space system, and experimental results show that our method achieves up to 98.35% system disruption against existing defense mechanisms while requiring significantly fewer feature modifications.
Yutong Cai, Ziang He, Yanyan Luo et al.· International Conference on...· 0 citations
: Massive multiple-input multiple-output (MIMO) technology is a key enabler for 5G and beyond wireless networks, offering significant improvements in spectral efficiency and link reliability. However, conventional beamforming techniques such as Zero Forcing (ZF) and Minimum Mean Square Error (MMSE) require complex matrix computations and fail to adapt efficiently to dynamic channel variations. To address these challenges, this paper proposes a Deep Deterministic Policy Gradient (DDPG)-based beamforming framework that formulates beamforming optimization as a continuous-action deep reinforcement learning problem. The proposed model directly generates complex-valued beamforming weight vectors to jointly maximize spectral efficiency (SE) and energy efficiency (EE) while minimizing the bit error rate (BER) and decision latency. An adaptive state representation incorporating channel state information (CSI), previous beamforming vectors, and performance metrics enables real-time policy learning under time-varying channel conditions. Simulation results demonstrate that the proposed method outperforms conventional and heuristic beamforming schemes, achieving up to 25% higher SE, 45% lower BER, 20% improvement in EE, and 30% latency reduction. The results validate the effectiveness of the proposed framework for energy-efficient, low-latency beamforming in next-generation massive MIMO wireless networks.
Nilakshee Rajule, Mithra Venkatesan, Harshada Magar et al.· Proceedings of the 1st Inter...· 0 citations
This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.
Mugerwa Joseph, Ajaegbu Chigozirim· International Journal Of Eng...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.