Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 41 references
TL;DR
The Dual-Critic Constrained Deceptive Q-Learning (DCD-Q) method is proposed, a deployment-time trajectory protection framework that aims to reduce the information leaked by released trajectories while preserving acceptable task performance.
Abstract
Securing the decision-making process of reinforcement learning (RL) agents during deployment is crucial in privacy-sensitive and security-critical domains. However, as deployed policies interact with the environment, they generate observable trajectories that can inadvertently leak sensitive decision patterns. External observers can readily exploit these trajectories via trajectory-based attacks such as imitation learning (IL), inverse reinforcement learning (IRL) to extract the policy or infer the underlying reward structure. Current deployment-time defenses typically tackle policy extraction or reward inference in isolation, and often fail to provide strict guarantees on the agent’s task utility. To bridge this gap, we propose the Dual-Critic Constrained Deceptive Q-Learning (DCD-Q) method, a deployment-time trajectory protection framework that aims to reduce the information leaked by released trajectories while preserving acceptable task performance. DCD-Q employs a utility critic to constrain decisions to a near-optimal candidate action set, aiming to preserve task utility while enabling controlled deceptive behavior. Simultaneously, it constructs a dynamic anti-reward from sliding-window visitation statistics to bias execution toward less-visited feasible actions. This induces a controlled, time-varying deceptive behavior in the released trajectories. We theoretically analyze the utility-preserving component by deriving a lower bound on the expected return under constrained action selection. Experiments on benchmark environments evaluate DCD-Q against representative trajectory-based attackers, which show that DCD-Q reduces the effectiveness of policy and reward recovery while maintaining a task-dependent utility–protection trade-off.
Client selection is a critical mechanism for ensuring robust convergence in Federated Learning (FL) systems, yet it remains vulnerable to Non-IID data distributions and Byzantine attacks. Deep Reinforcement Learning (DRL) has shown promise for automated client selection, yet existing methods suffer from three structura...
Empirical evaluations reveal a fundamental brittleness in existing defenses: with a single trainable 7B planner, Trident reduces blue agent defensive performance by an average of 522% compared to static red agent baselines while autonomously discovering emergent behaviors such as decoy avoidance and adaptive state prio...
Ryozo Masukawa, Ian Bryant, Armita Kazeminajafabadi et al.· 1 citation
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploita...
Rui-Han Wu, Rui Yang, Donggeon David Oh et al.· 0 citations
Reinforcement learning with verifiable rewards is typically performed on-policy, keeping training data close to the current policy but limiting learning to trajectories that the policy can discover itself. Off-policy methods such as supervised fine-tuning, on the other hand, can leverage external knowledge beyond the b...
Ming-Yu Chen, Ye-Fan Tao, Gerald Friedland et al.· 0 citations
Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy para...
Ali Irshayyid, Feng Lin, Chong Li et al.· 0 citations
A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...
A. Choudhury, J. Brahmanage, Akshat Kumar et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.