PoEM: Predicting RL Outcomes from Existing Policies
PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards, is introduced by introducing PoEM, a framework to predict the outputs of RL on a new reward function using a set of models already post-trained on other rewards.
K. Hamidieh, G. Daras, Antonio Torralba
· 0 citations