Open access
Aug 2026
Dual-critic constrained deceptive Q-learning for deployment-time policy protection
The Dual-Critic Constrained Deceptive Q-Learning (DCD-Q) method is proposed, a deployment-time trajectory protection framework that aims to reduce the information leaked by released trajectories while preserving acceptable task performance.
Guang-Yu Pan, Bo Hou, Yao Chen et al.
· Journal of King Saud Univers... · 0 citations