Hindsight-Anchored Policy Optimization: Learning Through Hindsight with Thompson Sampling-Inspired Adaptive Gating
HAPO improves the average accuracy of three major mixed-policy methods while maintaining training stability, and can be layered on top of existing mixed-policy methods in a generalizable manner.
Yuning Wu, Ke Wang, Hao Liu et al.
· 0 citations