Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning
We study federated reinforcement learning in which multiple agents interact with a common Markov decision process and communicate through a central server to collaboratively learn the optimal stateaction value function. Our goal is to understand whether the sample-efficiency benefits of collaboration can be retained wh...