Enhancing Stable Behavioral Imitation through Adaptive Reward Weighting in TD3-SAC-GAIL
The results demonstrate the potential of adaptive reward weighting to provide a systematic mechanism for controlling the exploration–imitation trade-off and enhancing the stability and robustness of GAIL-based policy learning while retaining the exploration advantages of the TD3-SAC hybrid framework.