This article addresses the multinode cooperative jamming problem in communication networks with unknown topology. To overcome the combinatorial explosion in decision-making and the credit assignment challenge under full-bandit feedback, we propose a novel reinforcement learning (RL) algorithm based on the multiarmed bandit (MAB) framework. The core of our approach is a collective-to-individual reward allocation model, which introduces a jamming overlap coefficient to quantify node interdependencies and employs the least absolute shrinkage and selection operator (LASSO) regression to estimate individual node contributions from the observed composite rewards. Building upon this, we design a triphase "collection-construction-exploitation" learning strategy. This strategy efficiently balances exploration and exploitation, enabling the jammers to progressively identify and focus on the most disruptive node combinations without any prior knowledge of the network structure. Theoretical analysis demonstrates that the algorithm achieves a sublinear regret bound under certain conditions. Comprehensive simulations across diverse network topologies and scenarios demonstrate that the proposed method significantly outperforms existing benchmarks in terms of cumulative regret, confirming its strong robustness and adaptability.
Cheng Zhou, Xin Man, Congshan Ma et al.· IEEE Transactions on Neural...· 0 citations
As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions that generalize across multiple environments. Because demonstrations in one MDP entangle reward information with that environments specific structure, the resulting rewards frequently fail to generalize when the agent is deployed in a new setting. We first analyze how different feedback modalities constrain rewards, showing that, in the unlimited-data regime, comparisons impose strictly stronger global constraints than other modalities. Beyond this theoretical analysis, we introduce a hierarchical machine teaching algorithm for reward learning that operates across multiple MDPs. The algorithm first greedily selects informative environments that expose complementary reward constraints, then strategically queries low-cost feedback within those environments. Empirically, our method achieves substantially lower regret and stronger generalization to held-out environments than uniform teaching baselines under identical feedback budgets, demonstrating the importance of multi-environment, multi-modal teaching for learning dynamics-robust reward functions.
Ali Larian, Qian Lin, Zong-Wu Chang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.